paper-with-me

Papers

A Fine-Grained Annotated Multi-Dialectal Arabic Corpus

2019-09-01 · RANLP 2019 9 · Anis Charfi, Wajdi Zaghouani, Syed Hassan Mehdi, Esraa Mohamed

We present ARAP-Tweet 2.0, a corpus of 5 million dialectal Arabic tweets and 50 million words of about 3000 Twitter users from 17 Arab countries. Compared to the first version, the new corpus has significant improvements in terms of the data volume and the annotation quality. It is fully balanced with respect to dialect, gender, and three age groups: under 25 years, between 25 and 34, and 35 years and above. This paper describes the process of creating the corpus starting from gathering the dialectal phrases to find the users, to annotating their accounts and retrieving their tweets. We also report on the evaluation of the annotation quality using the inter-annotator agreement measures which were applied to the whole corpus and not just a subset. The obtained results were substantial with average Cohen{'}s Kappa values of 0.99, 0.92, and 0.88 for the annotation of gender, dialect, and age respectively. We also discuss some challenges encountered when developing this corpus.s.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Multi-Dialect, Multi-Genre Corpus of Informal Written Arabic

2014-05-01 · LREC 2014 5 · Ryan Cotterell, Chris Callison-Burch

This paper presents a multi-dialect, multi-genre, human annotated corpus of dialectal Arabic. We collected utterances in five Arabic dialects: Levantine, Gulf, Egyptian, Iraqi and Maghrebi. We scraped newspaper websites …

Dialect Identification

ARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect Tagging

2026-01-05 · Omer Nacar, Serry Sibaee, Adel Ammar, Yasser Alhabashi 외 arxiv

The Arabic language is characterized by a rich tapestry of regional dialects that differ substantially in phonetics and lexicon, reflecting the geographic and cultural diversity of its speakers. Despite the availability …

Multi-Task Learning

Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues

2026-04-30 · Muhammad Dehan Al Kautsar, Saeed Almheiri, Momina Ahsan, Bilal Elbouardi 외 arxiv

There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standar…

Machine Translation

Munsit at NADI 2025 Shared Task 2: Pushing the Boundaries of Multidialectal Arabic ASR with Weakly Supervised Pretraining and Continual Supervised Fine-tuning

2025-08-12 · Mahmoud Salhab, Shameed Sait, Mohammad Abusheikh, Hasan Abusheikh arxiv

Automatic speech recognition (ASR) plays a vital role in enabling natural human-machine interaction across applications such as virtual assistants, industrial automation, customer support, and real-time transcription. Ho…

Speech Recognition

AraBench: Benchmarking Dialectal Arabic-English Machine Translation

2020-12-01 · COLING 2020 8 · Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, Fahim Dalvi

Low-resource machine translation suffers from the scarcity of training data and the unavailability of standard evaluation sets. While a number of research efforts target the former, the unavailability of evaluation bench…

BenchmarkingData AugmentationMachine TranslationTranslation