paper-with-me

홈 › Papers

Learning from Relatives: Unified Dialectal Arabic Segmentation

2017-08-01 · CONLL 2017 8 · Younes Samih, Mohamed Eldesouki, Mohammed Attia, Kareem Darwish, Ahmed Abdelali, Hamdy Mubarak, Laura Kallmeyer

Arabic dialects do not just share a common koin{\'e}, but there are shared pan-dialectal linguistic phenomena that allow computational models for dialects to learn from each other. In this paper we build a unified segmentation model where the training data for different dialects are combined and a single model is trained. The model yields higher accuracies than dialect-specific models, eliminating the need for dialect identification before segmentation. We also measure the degree of relatedness between four major Arabic dialects by testing how a segmentation model trained on one dialect performs on the other dialects. We found that linguistic relatedness is contingent with geographical proximity. In our experiments we use SVM-based ranking and bi-LSTM-CRF sequence labeling.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dialect IdentificationInformation RetrievalMachine TranslationSegmentation

Similar Papers 제목 키워드 기반

Morphology-aware Word-Segmentation in Dialectal Arabic Adaptation of Neural Machine Translation

2019-08-01 · WS 2019 8 · Ahmed Tawfik, Mahitab Emam, Khaled Essam, Robert Nabil 외

Parallel corpora available for building machine translation (MT) models for dialectal Arabic (DA) are rather limited. The scarcity of resources has prompted the use of Modern Standard Arabic (MSA) abundant resources to c…

Machine TranslationSegmentationTranslation

DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models

2025-10-31 · Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih 외 arxiv

We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluat…

Hierarchical Aggregation of Dialectal Data for Arabic Dialect Identification

2022-06-01 · LREC 2022 6 · Nurpeiis Baimukan, Houda Bouamor, Nizar Habash

Arabic is a collection of dialectal variants that are historically related but significantly different. These differences can be seen across regions, countries, and even cities in the same countries. Previous work on Ara…

Dialect Identification

Linear Semantic Segmentation for Low-Resource Spoken Dialects

2026-05-07 · Kirill Chirkunov, Younes Samih, Abed Alhakim Freihat, Hanan Aldarmaki arxiv

Semantic segmentation is a core component of discourse analysis, yet existing models are primarily developed and evaluated on high-resource written text, limiting their effectiveness on low-resource spoken varieties. In …

Semantic Segmentation

Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis

2026-01-20 · Yushen Chen, Junzhe Liu, Yujie Tu, Zhikang Niu 외 arxiv

Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. Key barriers include substantial cross-dialect lexical and phonological divergence, scarce synthesis-grade data, and the absen…

Speech Synthesis