paper-with-me

Papers

Building Dialectal Arabic Corpora

2017-09-01 · RANLP 2017 9 · Hani Elgabou, Dimitar Kazakov

The aim of this research is to identify local Arabic dialects in texts from social media (Twitter) and link them to specific geographic areas. Dialect identification is studied as a subset of the task of language identification. The proposed method is based on unsupervised learning using simultaneously lexical and geographic distance. While this study focusses on Libyan dialects, the approach is general, and could produce resources to support human translators and interpreters when dealing with vernaculars rather than standard Arabic.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dialect IdentificationInformation RetrievalLanguage IdentificationMachine Translation

Similar Papers 제목 키워드 기반

Morphology-aware Word-Segmentation in Dialectal Arabic Adaptation of Neural Machine Translation

2019-08-01 · WS 2019 8 · Ahmed Tawfik, Mahitab Emam, Khaled Essam, Robert Nabil 외

Parallel corpora available for building machine translation (MT) models for dialectal Arabic (DA) are rather limited. The scarcity of resources has prompted the use of Modern Standard Arabic (MSA) abundant resources to c…

Machine TranslationSegmentationTranslation

YADAC: Yet another Dialectal Arabic Corpus

2012-05-01 · LREC 2012 5 · Rania Al-Sabbagh, Roxana Girju

This paper presents the first phase of building YADAC ― a multi-genre Dialectal Arabic (DA) corpus ― that is compiled using Web data from microblogs (i.e. Twitter), blogs/forums and online knowledge market services in wh…

ChunkingDialect IdentificationPart-Of-Speech TaggingPOS+2

Toward a Web-based Speech Corpus for Algerian Dialectal Arabic Varieties

2017-04-01 · WS 2017 4 · Soumia Bougrine, Aicha Chorana, Abdallah Lakhdari, Hadda Cherroun

The success of machine learning for automatic speech processing has raised the need for large scale datasets. However, collecting such data is often a challenging task as it implies significant investment involving time …

Speech RecognitionSpeech Synthesis

Hierarchical Aggregation of Dialectal Data for Arabic Dialect Identification

2022-06-01 · LREC 2022 6 · Nurpeiis Baimukan, Houda Bouamor, Nizar Habash

Arabic is a collection of dialectal variants that are historically related but significantly different. These differences can be seen across regions, countries, and even cities in the same countries. Previous work on Ara…

Dialect Identification

DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models

2025-10-31 · Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih 외 arxiv

We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluat…