paper-with-me

홈 › Papers

Processing Dialectal Arabic: Exploiting Variability and Similarity to Overcome Challenges and Discover Opportunities

2016-12-01 · WS 2016 12 · Mona Diab

We recently witnessed an exponential growth in dialectal Arabic usage in both textual data and speech recordings especially in social media. Processing such media is of great utility for all kinds of applications ranging from information extraction to social media analytics for political and commercial purposes to building decision support systems. Compared to other languages, Arabic, especially the informal variety, poses a significant challenge to natural language processing algorithms since it comprises multiple dialects, linguistic code switching, and a lack of standardized orthographies, to top its relatively complex morphology. Inherently, the problem of processing Arabic in the context of social media is the problem of how to handle resource poor languages. In this talk I will go over some of our insights to some of these problems and show how there is a silver lining where we can generalize some of our solutions to other low resource language contexts.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Aladdin-FTI @ AMIYA Three Wishes for Arabic NLP: Fidelity, Diglossia, and Multidialectal Generation

2026-02-18 · Jonathan Mutal, Perla Al Almaoui, Simon Hengchen, Pierrette Bouillon arxiv

Arabic dialects have long been under-represented in Natural Language Processing (NLP) research due to their non-standardization and high variability, which pose challenges for computational modeling. Recent advances in t…

Text Generation

Exploiting Dialect Identification in Automatic Dialectal Text Normalization

2024-07-03 · Bashar Alhafni, Sarah Al-Towaity, Ziyad Fawzy, Fatema Nassar 외

Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do no…

Dialect IdentificationText Normalization

Casablanca: Data and Models for Multidialectal Arabic Speech Recognition

2024-10-06 · Bashar Talafha, Karima Kadaoui, Samar Mohamed Magdy, Mariem Habiboullah 외

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and …

Arabic Speech Recognitionspeech-recognitionSpeech Recognition

ELYADATA & LIA at NADI 2025: ASR and ADI Subtasks

2025-11-13 · Haroun Elleuch, Youssef Saidi, Salima Mdhaffar, Yannick Estève 외 arxiv

This paper describes Elyadata \& LIA's joint submission to the NADI multi-dialectal Arabic Speech Processing 2025. We participated in the Spoken Arabic Dialect Identification (ADI) and multi-dialectal Arabic ASR subtasks…

Data Augmentation

Creating Resources for Dialectal Arabic from a Single Annotation: A Case Study on Egyptian and Levantine

2016-12-01 · COLING 2016 12 · Esk, Ramy er, Nizar Habash, Owen Rambow 외

Arabic dialects present a special problem for natural language processing because there are few resources, they have no standard orthography, and have not been studied much. However, as more and more written dialectal Ar…

Morphological Analysis