Identification of Languages in Algerian Arabic Multilingual Documents
This paper presents a language identification system designed to detect the language of each word, in its context, in a multilingual documents as generated in social media by bilingual/multilingual communities, in our case speakers of Algerian Arabic. We frame the task as a sequence tagging problem and use supervised machine learning with standard methods like HMM and Ngram classification tagging. We also experiment with a lexicon-based method. Combining all the methods in a fall-back mechanism and introducing some linguistic rules, to deal with unseen tokens and ambiguous words, gives an overall accuracy of 93.14{\%}. Finally, we introduced rules for language identification from sequences of recognised words.
Code (0)
등록된 구현이 없습니다.
Tasks
ChunkingGeneral ClassificationLanguage IdentificationSimilar Papers 제목 키워드 기반
DziriBERT: a Pre-trained Language Model for the Algerian Dialect
Pre-trained transformers are now the de facto models in Natural Language Processing given their state-of-the-art results in many tasks and languages. However, most of the current models have been trained on languages for…
Language ModelingLanguage ModellingHate speech detection in algerian dialect using deep learning
With the proliferation of hate speech on social networks under different formats, such as abusive language, cyberbullying, and violence, etc., people have experienced a significant increase in violence, putting them in u…
Abusive LanguageDeep LearningHate Speech DetectionOffensive Language Detection in Under-resourced Algerian Dialectal Arabic Language
This paper addresses the problem of detecting the offensive and abusive content in Facebook comments, where we focus on the Algerian dialectal Arabic which is one of under-resourced languages. The latter has a variety of…
Hierarchical Classification for Spoken Arabic Dialect Identification using Prosody: Case of Algerian Dialects
In daily communications, Arabs use local dialects which are hard to identify automatically using conventional classification methods. The dialect identification challenging task becomes more complicated when dealing with…
ClassificationDialect IdentificationGeneral ClassificationA Comparison of Character Neural Language Model and Bootstrapping for Language Identification in Multilingual Noisy Texts
This paper seeks to examine the effect of including background knowledge in the form of character pre-trained neural language model (LM), and data bootstrapping to overcome the problem of unbalanced limited resources. As…
Language IdentificationLanguage ModelingLanguage ModellingMulti-Task Learning