paper-with-me

홈 › Papers

ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks

2025-11-13 · Haroun Elleuch, Youssef Saidi, Salima Mdhaffar, Yannick Estève, Fethi Bougares arxiv

This paper describes Elyadata \& LIA's joint submission to the NADI multi-dialectal Arabic Speech Processing 2025. We participated in the Spoken Arabic Dialect Identification (ADI) and multi-dialectal Arabic ASR subtasks. Our submission ranked first for the ADI subtask and second for the multi-dialectal Arabic ASR subtask among all participants. Our ADI system is a fine-tuned Whisper-large-v3 encoder with data augmentation. This system obtained the highest ADI accuracy score of \textbf{79.83\%} on the official test set. For multi-dialectal Arabic ASR, we fine-tuned SeamlessM4T-v2 Large (Egyptian variant) separately for each of the eight considered dialects. Overall, we obtained an average WER and CER of \textbf{38.54\%} and \textbf{14.53\%}, respectively, on the test set. Our results demonstrate the effectiveness of large pre-trained speech models with targeted fine-tuning for Arabic speech processing.

📄 PDF Abstract BibTeX arXiv:2511.10090

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

NADI 2023: The Fourth Nuanced Arabic Dialect Identification Shared Task

2023-10-24 · Muhammad Abdul-Mageed, AbdelRahim Elmadany, Chiyu Zhang, El Moatez Billah Nagoudi 외

We describe the findings of the fourth Nuanced Arabic Dialect Identification Shared Task (NADI 2023). The objective of NADI is to help advance state-of-the-art Arabic NLP by creating opportunities for teams of researcher…

Dialect IdentificationMachine Translationvalid

NADI 2022: The Third Nuanced Arabic Dialect Identification Shared Task

2022-10-18 · Muhammad Abdul-Mageed, Chiyu Zhang, AbdelRahim Elmadany, Houda Bouamor 외

We describe findings of the third Nuanced Arabic Dialect Identification Shared Task (NADI 2022). NADI aims at advancing state of the art Arabic NLP, including on Arabic dialects. It does so by affording diverse datasets …

Dialect IdentificationSentiment Analysisvalid

NADI 2020: The First Nuanced Arabic Dialect Identification Shared Task

2020-10-21 · COLING (WANLP) 2020 12 · Muhammad Abdul-Mageed, Chiyu Zhang, Houda Bouamor, Nizar Habash

We present the results and findings of the First Nuanced Arabic Dialect Identification Shared Task (NADI). This Shared Task includes two subtasks: country-level dialect identification (Subtask 1) and province-level sub-d…

Dialect Identification

Adapting MARBERT for Improved Arabic Dialect Identification: Submission to the NADI 2021 Shared Task

2021-03-01 · EACL (WANLP) 2021 4 · Badr AlKhamissi, Mohamed Gabr, Muhammad ElNokrashy, Khaled Essam

In this paper, we tackle the Nuanced Arabic Dialect Identification (NADI) shared task (Abdul-Mageed et al., 2021) and demonstrate state-of-the-art results on all of its four subtasks. Tasks are to identify the geographic…

Dialect Identification

Machine Learning-Based Approach for Arabic Dialect Identification

2021-04-01 · EACL (WANLP) 2021 4 · Hamada Nayel, Ahmed Hassan, Mahmoud Sobhi, Ahmed El-Sawy

This paper describes our systems submitted to the Second Nuanced Arabic Dialect Identification Shared Task (NADI 2021). Dialect identification is the task of automatically detecting the source variety of a given text or …

BIG-bench Machine LearningDialect Identificationregression