paper-with-me

홈 › Papers

ALDi: Quantifying the Arabic Level of Dialectness of Text

2023-10-20 · Amr Keleg, Sharon Goldwater, Walid Magdy

Transcribed speech and user-generated text in Arabic typically contain a mixture of Modern Standard Arabic (MSA), the standardized language taught in schools, and Dialectal Arabic (DA), used in daily communications. To handle this variation, previous work in Arabic NLP has focused on Dialect Identification (DI) on the sentence or the token level. However, DI treats the task as binary, whereas we argue that Arabic speakers perceive a spectrum of dialectness, which we operationalize at the sentence level as the Arabic Level of Dialectness (ALDi), a continuous linguistic variable. We introduce the AOC-ALDi dataset (derived from the AOC dataset), containing 127,835 sentences (17% from news articles and 83% from user comments on those articles) which are manually labeled with their level of dialectness. We provide a detailed analysis of AOC-ALDi and show that a model trained on it can effectively identify levels of dialectness on a range of other corpora (including dialects and genres not included in AOC-ALDi), providing a more nuanced picture than traditional DI systems. Through case studies, we illustrate how ALDi can reveal Arabic speakers' stylistic choices in different situations, a useful property for sociolinguistic analyses.

📄 PDF Abstract BibTeX arXiv:2310.13747

Code (1)

amr-keleg/aldi 공식 구현

Tasks

ArticlesDialect IdentificationSentence

Similar Papers 제목 키워드 기반

The Arabic Generality Score: Another Dimension of Modeling Arabic Dialectness

2025-08-24 · Sanad Shaban, Nizar Habash arxiv

Arabic dialects form a diverse continuum, yet NLP models often treat them as discrete categories. Recent work addresses this issue by modeling dialectness as a continuous variable, notably through the Arabic Level of Dia…

Word Alignment

Estimating the Level of Dialectness Predicts Interannotator Agreement in Multi-dialect Arabic Datasets

2024-05-18 · Amr Keleg, Walid Magdy, Sharon Goldwater

On annotating multi-dialect Arabic datasets, it is common to randomly assign the samples across a pool of native Arabic speakers. Recent analyses recommended routing dialectal samples to native speakers of their respecti…

SentenceSentence Classification

Curriculum Learning and Pseudo-Labeling Improve the Generalization of Multi-Label Arabic Dialect Identification Models

2026-02-12 · Ali Mekky, Mohamed El Zeftawy, Lara Hassan, Amr Keleg 외 arxiv

Being modeled as a single-label classification task for a long time, recent work has argued that Arabic Dialect Identification (ADI) should be framed as a multi-label classification task. However, ADI remains constrained…

Multi-Label Classification

AraBench: Benchmarking Dialectal Arabic-English Machine Translation

2020-12-01 · COLING 2020 8 · Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, Fahim Dalvi

Low-resource machine translation suffers from the scarcity of training data and the unavailability of standard evaluation sets. While a number of research efforts target the former, the unavailability of evaluation bench…

BenchmarkingData AugmentationMachine TranslationTranslation

Morphotactic Modeling in an Open-source Multi-dialectal Arabic Morphological Analyzer and Generator

2022-07-01 · NAACL (SIGMORPHON) 2022 7 · Nizar Habash, Reham Marzouk, Christian Khairallah, Salam Khalifa

Arabic is a morphologically rich and complex language, with numerous dialectal variants. Previous efforts on Arabic morphology modeling focused on specific variants and specific domains using a range of techniques with d…