Romanized Berber and Romanized Arabic Automatic Language Identification Using Machine Learning
The identification of the language of text/speech input is the first step to be able to properly do any language-dependent natural language processing. The task is called Automatic Language Identification (ALI). Being a well-studied field since early 1960{'}s, various methods have been applied to many standard languages. The ALI standard methods require datasets for training and use character/word-based n-gram models. However, social media and new technologies have contributed to the rise of informal and minority languages on the Web. The state-of-the-art automatic language identifiers fail to properly identify many of them. Romanized Arabic (RA) and Romanized Berber (RB) are cases of these informal languages which are under-resourced. The goal of this paper is twofold: detect RA and RB, at a document level, as separate languages and distinguish between them as they coexist in North Africa. We consider the task as a classification problem and use supervised machine learning to solve it. For both languages, character-based 5-grams combined with additional lexicons score the best, F-score of 99.75{\%} and 97.77{\%} for RB and RA respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningLanguage IdentificationTransliterationSimilar Papers 제목 키워드 기반
Automatic Transliteration of Romanized Dialectal Arabic
Finding Romanized Arabic Dialect in Code-Mixed Tweets
Recent computational work on Arabic dialect identification has focused primarily on building and annotating corpora written in Arabic script. Arabic dialects however also appear written in Roman script, especially in soc…
Dialect IdentificationLanguage IdentificationRomanized Arabic Transliteration
Phonetic and Visual Priors for Decipherment of Informal Romanization
Informal romanization is an idiosyncratic process used by humans in informal digital communication to encode non-Latin script languages into Latin character sets found on common keyboards. Character substitution choices …
DeciphermentInductive BiasAutomatic Romanization of Arabic Bibliographic Records
International library standards require cataloguers to tediously input Romanization of their catalogue records for the benefit of library users without specific language expertise. In this paper, we present the first rep…