Machine Learning-Based Approach for Arabic Dialect Identification
This paper describes our systems submitted to the Second Nuanced Arabic Dialect Identification Shared Task (NADI 2021). Dialect identification is the task of automatically detecting the source variety of a given text or speech segment. There are four subtasks, two subtasks for country-level identification and the other two subtasks for province-level identification. The data in this task covers a total of 100 provinces from all 21 Arab countries and come from the Twitter domain. The proposed systems depend on five machine-learning approaches namely Complement Naïve Bayes, Support Vector Machine, Decision Tree, Logistic Regression and Random Forest Classifiers. F1 macro-averaged score of Naïve Bayes classifier outperformed all other classifiers for development and test data.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningDialect IdentificationregressionSimilar Papers 제목 키워드 기반
Automatic Arabic Dialect Identification Systems for Written Texts: A Survey
Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural lang…
Dialect IdentificationMachine TranslationSentenceSpeech Synthesis+6Comparing Pipelined and Integrated Approaches to Dialectal Arabic Neural Machine Translation
When translating diglossic languages such as Arabic, situations may arise where we would like to translate a text but do not know which dialect it is. A traditional approach to this problem is to design dialect identific…
Dialect IdentificationMachine TranslationTranslationNADI 2024: The Fifth Nuanced Arabic Dialect Identification Shared Task
We describe the findings of the fifth Nuanced Arabic Dialect Identification Shared Task (NADI 2024). NADI's objective is to help advance SoTA Arabic NLP by providing guidance, datasets, modeling opportunities, and standa…
Dialect IdentificationMachine TranslationTranslationvalidAutomatic Dialect Detection in Arabic Broadcast Speech
We investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. W…
Dialect IdentificationLanguage Identificationspeech-recognitionSpeech Recognition+1Weighted combination of BERT and N-GRAM features for Nuanced Arabic Dialect Identification
Around the Arab world, different Arabic dialects are spoken by more than 300M persons, and are increasingly popular in social media texts. However, Arabic dialects are considered to be low-resource languages, limiting th…
Dialect Identification