paper-with-me

Papers

BERT-based Multi-Task Model for Country and Province Level Modern Standard Arabic and Dialectal Arabic Identification

2021-06-23 · Abdellah El Mekki, Abdelkader El Mahdaouy, Kabil Essefar, Nabil El Mamoun, Ismail Berrada, Ahmed Khoumsi

Dialect and standard language identification are crucial tasks for many Arabic natural language processing applications. In this paper, we present our deep learning-based system, submitted to the second NADI shared task for country-level and province-level identification of Modern Standard Arabic (MSA) and Dialectal Arabic (DA). The system is based on an end-to-end deep Multi-Task Learning (MTL) model to tackle both country-level and province-level MSA/DA identification. The latter MTL model consists of a shared Bidirectional Encoder Representation Transformers (BERT) encoder, two task-specific attention layers, and two classifiers. Our key idea is to leverage both the task-discriminative and the inter-task shared features for country and province MSA/DA identification. The obtained results show that our MTL model outperforms single-task models on most subtasks.

📄 PDF Abstract BibTeX arXiv:2106.12495

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationMulti-Task Learning

Similar Papers 제목 키워드 기반

BERT-based Multi-Task Model for Country and Province Level MSA and Dialectal Arabic Identification

2021-04-01 · EACL (WANLP) 2021 4 · Abdellah El Mekki, Abdelkader El Mahdaouy, Kabil Essefar, Nabil El Mamoun 외

Dialect and standard language identification are crucial tasks for many Arabic natural language processing applications. In this paper, we present our deep learning-based system, submitted to the second NADI shared task …

Language IdentificationMulti-Task Learning

Weighted combination of BERT and N-GRAM features for Nuanced Arabic Dialect Identification

2020-12-01 · COLING (WANLP) 2020 12 · Abdellah El Mekki, Ahmed Alami, Hamza Alami, Ahmed Khoumsi 외

Around the Arab world, different Arabic dialects are spoken by more than 300M persons, and are increasingly popular in social media texts. However, Arabic dialects are considered to be low-resource languages, limiting th…

Dialect Identification

Adapting MARBERT for Improved Arabic Dialect Identification: Submission to the NADI 2021 Shared Task

2021-03-01 · EACL (WANLP) 2021 4 · Badr AlKhamissi, Mohamed Gabr, Muhammad ElNokrashy, Khaled Essam

In this paper, we tackle the Nuanced Arabic Dialect Identification (NADI) shared task (Abdul-Mageed et al., 2021) and demonstrate state-of-the-art results on all of its four subtasks. Tasks are to identify the geographic…

Dialect Identification

Arabic dialect identification: An Arabic-BERT model with data augmentation and ensembling strategy

2020-12-01 · COLING (WANLP) 2020 12 · Kamel Gaanoun, Imade Benelallam

This paper presents the ArabicProcessors team’s deep learning system designed for the NADI 2020 Subtask 1 (country-level dialect identification) and Subtask 2 (province-level dialect identification). We used Arabic-Bert …

Data AugmentationDialect Identification

NADI 2021: The Second Nuanced Arabic Dialect Identification Shared Task

2021-03-04 · EACL (WANLP) 2021 4 · Muhammad Abdul-Mageed, Chiyu Zhang, AbdelRahim Elmadany, Houda Bouamor 외

We present the findings and results of the Second Nuanced Arabic Dialect Identification Shared Task (NADI 2021). This Shared Task includes four subtasks: country-level Modern Standard Arabic (MSA) identification (Subtask…

Dialect Identification