paper-with-me

홈 › Papers

CODET: A Benchmark for Contrastive Dialectal Evaluation of Machine Translation

2023-05-26 · Md Mahfuz ibn Alam, Sina Ahmadi, Antonios Anastasopoulos

Neural machine translation (NMT) systems exhibit limited robustness in handling source-side linguistic variations. Their performance tends to degrade when faced with even slight deviations in language usage, such as different domains or variations introduced by second-language speakers. It is intuitive to extend this observation to encompass dialectal variations as well, but the work allowing the community to evaluate MT systems on this dimension is limited. To alleviate this issue, we compile and release CODET, a contrastive dialectal benchmark encompassing 891 different variations from twelve different languages. We also quantitatively demonstrate the challenges large MT models face in effectively translating dialectal variants. All the data and code have been released.

📄 PDF Abstract BibTeX arXiv:2305.17267

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

AraBench: Benchmarking Dialectal Arabic-English Machine Translation

2020-12-01 · COLING 2020 8 · Hassan Sajjad, Ahmed Abdelali, Nadir Durrani, Fahim Dalvi

Low-resource machine translation suffers from the scarcity of training data and the unavailability of standard evaluation sets. While a number of research efforts target the former, the unavailability of evaluation bench…

BenchmarkingData AugmentationMachine TranslationTranslation

CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation

2023-10-08 · Weixiang Yan, Yuchen Tian, Yunzhe Li, Qian Chen 외

Recent code translation techniques exploit neural machine translation models to translate source code from one programming language to another to satisfy production compatibility or to improve efficiency of codebase main…

Code TranslationMachine TranslationTranslation

AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs

2024-09-17 · Basel Mousi, Nadir Durrani, Fatema Ahmad, Md. Arid Hasan 외

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects …

Dialect IdentificationDiversityMachine TranslationTranslation

DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models

2025-10-31 · Malik H. Altakrori, Nizar Habash, Abed Alhakim Freihat, Younes Samih 외 arxiv

We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluat…

Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs

2026-01-19 · Abdellah El Mekki, Samar M. Magdy, Houdaifa Atou, Ruwa AbuHweidi 외 arxiv

Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than Modern Standard Arabic (MSA). Despite this, machine translation (MT) systems often generalize poorly to dialect…

Machine Translation