Country-level Arabic Dialect Identification using RNNs with and without Linguistic Features
This work investigates the value of augmenting recurrent neural networks with feature engineering for the Second Nuanced Arabic Dialect Identification (NADI) Subtask 1.2: Country-level DA identification. We compare the performance of a simple word-level LSTM using pretrained embeddings with one enhanced using feature embeddings for engineered linguistic features. Our results show that the addition of explicit features to the LSTM is detrimental to performance. We attribute this performance loss to the bivalency of some linguistic items in some text, ubiquity of topics, and participant mobility.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeDialect IdentificationFeature EngineeringSimilar Papers 제목 키워드 기반
Weighted combination of BERT and N-GRAM features for Nuanced Arabic Dialect Identification
Around the Arab world, different Arabic dialects are spoken by more than 300M persons, and are increasingly popular in social media texts. However, Arabic dialects are considered to be low-resource languages, limiting th…
Dialect IdentificationBERT-based Multi-Task Model for Country and Province Level Modern Standard Arabic and Dialectal Arabic Identification
Dialect and standard language identification are crucial tasks for many Arabic natural language processing applications. In this paper, we present our deep learning-based system, submitted to the second NADI shared task …
Language IdentificationMulti-Task LearningBERT-based Multi-Task Model for Country and Province Level MSA and Dialectal Arabic Identification
Dialect and standard language identification are crucial tasks for many Arabic natural language processing applications. In this paper, we present our deep learning-based system, submitted to the second NADI shared task …
Language IdentificationMulti-Task LearningHierarchical Aggregation of Dialectal Data for Arabic Dialect Identification
Arabic is a collection of dialectal variants that are historically related but significantly different. These differences can be seen across regions, countries, and even cities in the same countries. Previous work on Ara…
Dialect IdentificationNADI 2021: The Second Nuanced Arabic Dialect Identification Shared Task
We present the findings and results of the Second Nuanced Arabic Dialect Identification Shared Task (NADI 2021). This Shared Task includes four subtasks: country-level Modern Standard Arabic (MSA) identification (Subtask…
Dialect Identification