paper-with-me

홈 › Papers

Exploiting Out-of-Domain Data Sources for Dialectal Arabic Statistical Machine Translation

2015-09-07 · Katrin Kirchhoff, Bing Zhao, Wen Wang

Statistical machine translation for dialectal Arabic is characterized by a lack of data since data acquisition involves the transcription and translation of spoken language. In this study we develop techniques for extracting parallel data for one particular dialect of Arabic (Iraqi Arabic) from out-of-domain corpora in different dialects of Arabic or in Modern Standard Arabic. We compare two different data selection strategies (cross-entropy based and submodular selection) and demonstrate that a very small but highly targeted amount of found data can improve the performance of a baseline machine translation system. We furthermore report on preliminary experiments on using automatically translated speech data as additional training data.

📄 PDF Abstract BibTeX arXiv:1509.01938

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Content-Localization based Neural Machine Translation for Informal Dialectal Arabic: Spanish/French to Levantine/Gulf Arabic

2023-12-12 · Fatimah Alzamzami, Abdulmotaleb El Saddik

Resources in high-resource languages have not been efficiently exploited in low-resource languages to solve language-dependent research problems. Spanish and French are considered high resource languages in which an adeq…

Machine TranslationTranslation

Exploiting Dialect Identification in Automatic Dialectal Text Normalization

2024-07-03 · Bashar Alhafni, Sarah Al-Towaity, Ziyad Fawzy, Fatema Nassar 외

Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do no…

Dialect IdentificationText Normalization

Creating Resources for Dialectal Arabic from a Single Annotation: A Case Study on Egyptian and Levantine

2016-12-01 · COLING 2016 12 · Esk, Ramy er, Nizar Habash, Owen Rambow 외

Arabic dialects present a special problem for natural language processing because there are few resources, they have no standard orthography, and have not been studied much. However, as more and more written dialectal Ar…

Morphological Analysis

Morphology-aware Word-Segmentation in Dialectal Arabic Adaptation of Neural Machine Translation

2019-08-01 · WS 2019 8 · Ahmed Tawfik, Mahitab Emam, Khaled Essam, Robert Nabil 외

Parallel corpora available for building machine translation (MT) models for dialectal Arabic (DA) are rather limited. The scarcity of resources has prompted the use of Modern Standard Arabic (MSA) abundant resources to c…

Machine TranslationSegmentationTranslation

LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models

2026-02-17 · Ahmed Khaled Khamis, Hesham Ali arxiv

Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic …

Synthetic Data GenerationSpeaker DiarizationSpeech SynthesisText to Speech