ArbDialectID at MADAR Shared Task 1: Language Modelling and Ensemble Learning for Fine Grained Arabic Dialect Identification
In this paper, we present a Dialect Identification system (ArbDialectID) that competed at Task 1 of the MADAR shared task, MADARTravel Domain Dialect Identification. We build a course and a fine-grained identification model to predict the label (corresponding to a dialect of Arabic) of a given text. We build two language models by extracting features at two levels (words and characters). We firstly build a coarse identification model to classify each sentence into one out of six dialects, then use this label as a feature for the fine-grained model that classifies the sentence among 26 dialects from different Arab cities, after that we apply ensemble voting classifier on both sub-systems. Our system ranked 1st that achieving an f-score of 67.32{\%}. Both the models and our feature engineering tools are made available to the research community.
Code (0)
등록된 구현이 없습니다.
Tasks
Dialect IdentificationEnsemble LearningFeature EngineeringLanguage ModellingSentenceSimilar Papers 제목 키워드 기반
ZCU-NLP at MADAR 2019: Recognizing Arabic Dialects
In this paper, we present our systems for the MADAR Shared Task: Arabic Fine-Grained Dialect Identification. The shared task consists of two subtasks. The goal of Subtask{--} 1 (S-1) is to detect an Arabic city dialect i…
BIG-bench Machine LearningDialect IdentificationLanguage ModellingThe MADAR Shared Task on Arabic Fine-Grained Dialect Identification
In this paper, we present the results and findings of the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. This shared task was organized as part of The Fourth Arabic Natural Language Processing Workshop,…
Dialect IdentificationLIUM-MIRACL Participation in the MADAR Arabic Dialect Identification Shared Task
This paper describes the joint participation of the LIUM and MIRACL Laboratories at the Arabic dialect identification challenge of the MADAR Shared Task (Bouamor et al., 2019) conducted during the Fourth Arabic Natural L…
Deep LearningDialect IdentificationSentenceJHU System Description for the MADAR Arabic Dialect Identification Shared Task
Our submission to the MADAR shared task on Arabic dialect identification employed a language modeling technique called Prediction by Partial Matching, an ensemble of neural architectures, and sources of additional data f…
Dialect IdentificationLanguage ModelingLanguage ModellingWord EmbeddingsTeam JUST at the MADAR Shared Task on Arabic Fine-Grained Dialect Identification
In this paper, we describe our team{'}s effort on the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. The task requires building a system capable of differentiating between 25 different Arabic dialects i…
Data AugmentationDialect IdentificationLanguage ModelingLanguage Modelling