paper-with-me

홈 › Papers

PALI: A Language Identification Benchmark for Perso-Arabic Scripts

2023-04-03 · Sina Ahmadi, Milind Agarwal, Antonios Anastasopoulos

The Perso-Arabic scripts are a family of scripts that are widely adopted and used by various linguistic communities around the globe. Identifying various languages using such scripts is crucial to language technologies and challenging in low-resource setups. As such, this paper sheds light on the challenges of detecting languages using Perso-Arabic scripts, especially in bilingual communities where ``unconventional'' writing is practiced. To address this, we use a set of supervised techniques to classify sentences into their languages. Building on these, we also propose a hierarchical model that targets clusters of languages that are more often confused by the classifiers. Our experiment results indicate the effectiveness of our solutions.

📄 PDF Abstract BibTeX arXiv:2304.01322

Code (1)

sinaahmadi/persoarabiclid 공식 구현

Tasks

Language Identification

Similar Papers 제목 키워드 기반

LILI: A Simple Language Independent Approach for Language Identification

2016-12-01 · COLING 2016 12 · Mohamed Al-Badrashiny, Mona Diab

We introduce a generic Language Independent Framework for Linguistic Code Switch Point Detection. The system uses characters level 5-grams and word level unigram language models to train a conditional random fields (CRF)…

Language Identification

LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation

2020-05-09 · LREC 2020 5 · Gustavo Aguilar, Sudipta Kar, Thamar Solorio

Recent trends in NLP research have raised an interest in linguistic code-switching (CS); modern approaches have been proposed to solve a wide range of NLP tasks on multiple language pairs. Unfortunately, these proposed m…

Language Identificationnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+2

Automatic Gender Identification and Reinflection in Arabic

2019-08-01 · WS 2019 8 · Nizar Habash, Houda Bouamor, Christine Chung

The impressive progress in many Natural Language Processing (NLP) applications has increased the awareness of some of the biases these NLP systems have with regards to gender identities. In this paper, we propose an appr…

Machine TranslationTranslation

The Arabic Parallel Gender Corpus 2.0: Extensions and Analyses

2021-10-18 · LREC 2022 6 · Bashar Alhafni, Nizar Habash, Houda Bouamor

Gender bias in natural language processing (NLP) applications, particularly machine translation, has been receiving increasing attention. Much of the research on this issue has focused on mitigating gender bias in Englis…

Machine TranslationText GenerationTranslation

Weighted combination of BERT and N-GRAM features for Nuanced Arabic Dialect Identification

2020-12-01 · COLING (WANLP) 2020 12 · Abdellah El Mekki, Ahmed Alami, Hamza Alami, Ahmed Khoumsi 외

Around the Arab world, different Arabic dialects are spoken by more than 300M persons, and are increasingly popular in social media texts. However, Arabic dialects are considered to be low-resource languages, limiting th…

Dialect Identification