Mawdoo3 AI at MADAR Shared Task: Arabic Tweet Dialect Identification
Arabic dialect identification is an inherently complex problem, as Arabic dialect taxonomy is convoluted and aims to dissect a continuous space rather than a discrete one. In this work, we present machine and deep learning approaches to predict 21 fine-grained dialects form a set of given tweets per user. We adopted numerous feature extraction methods most of which showed improvement in the final model, such as word embedding, Tf-idf, and other tweet features. Our results show that a simple LinearSVC can outperform any complex deep learning model given a set of curated features. With a relatively complex user voting mechanism, we were able to achieve a Macro-Averaged F1-score of 71.84{\%} on MADAR shared subtask-2. Our best submitted model ranked second out of all participating teams.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningDialect IdentificationSimilar Papers 제목 키워드 기반
Mawdoo3 AI at MADAR Shared Task: Arabic Fine-Grained Dialect Identification with Ensemble Learning
In this paper we discuss several models we used to classify 25 city-level Arabic dialects in addition to Modern Standard Arabic (MSA) as part of MADAR shared task (sub-task 1). We propose an ensemble model of a group of …
Dialect IdentificationEnsemble LearningZCU-NLP at MADAR 2019: Recognizing Arabic Dialects
In this paper, we present our systems for the MADAR Shared Task: Arabic Fine-Grained Dialect Identification. The shared task consists of two subtasks. The goal of Subtask{--} 1 (S-1) is to detect an Arabic city dialect i…
BIG-bench Machine LearningDialect IdentificationLanguage ModellingThe MADAR Shared Task on Arabic Fine-Grained Dialect Identification
In this paper, we present the results and findings of the MADAR Shared Task on Arabic Fine-Grained Dialect Identification. This shared task was organized as part of The Fourth Arabic Natural Language Processing Workshop,…
Dialect IdentificationMulti-Dialect Arabic BERT for Country-Level Dialect Identification
Arabic dialect identification is a complex problem for a number of inherent properties of the language itself. In this paper, we present the experiments conducted, and the models developed by our competing team, Mawdoo3 …
Dialect IdentificationLanguage ModelingLanguage ModellingNo Army, No Navy: BERT Semi-Supervised Learning of Arabic Dialects
We present our deep leaning system submitted to MADAR shared task 2 focused on twitter user dialect identification. We develop tweet-level identification models based on GRUs and BERT in supervised and semi-supervised se…
Dialect IdentificationTask 2