Building Dialectal Arabic Corpora
The aim of this research is to identify local Arabic dialects in texts from social media (Twitter) and link them to specific geographic areas. Dialect identification is studied as a subset of the task of language identification. The proposed method is based on unsupervised learning using simultaneously lexical and geographic distance. While this study focusses on Libyan dialects, the approach is general, and could produce resources to support human translators and interpreters when dealing with vernaculars rather than standard Arabic.
Code (0)
등록된 구현이 없습니다.
Tasks
Dialect IdentificationInformation RetrievalLanguage IdentificationMachine TranslationSimilar Papers 제목 키워드 기반
Morphology-aware Word-Segmentation in Dialectal Arabic Adaptation of Neural Machine Translation
Parallel corpora available for building machine translation (MT) models for dialectal Arabic (DA) are rather limited. The scarcity of resources has prompted the use of Modern Standard Arabic (MSA) abundant resources to c…
Machine TranslationSegmentationTranslationYADAC: Yet another Dialectal Arabic Corpus
This paper presents the first phase of building YADAC ― a multi-genre Dialectal Arabic (DA) corpus ― that is compiled using Web data from microblogs (i.e. Twitter), blogs/forums and online knowledge market services in wh…
ChunkingDialect IdentificationPart-Of-Speech TaggingPOS+2Toward a Web-based Speech Corpus for Algerian Dialectal Arabic Varieties
The success of machine learning for automatic speech processing has raised the need for large scale datasets. However, collecting such data is often a challenging task as it implies significant investment involving time …
Speech RecognitionSpeech SynthesisHierarchical Aggregation of Dialectal Data for Arabic Dialect Identification
Arabic is a collection of dialectal variants that are historically related but significantly different. These differences can be seen across regions, countries, and even cities in the same countries. Previous work on Ara…
Dialect IdentificationDialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluat…