paper-with-me

홈 › Papers

Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages

2026-03-26 · Danlu Chen, Ka Sing He, Jiahe Tian, Chenghao Xiao, Zhaofeng Wu, Taylor Berg-Kirkpatrick, Freda Shi arxiv

The landscape of extremely low-resource machine translation (MT) is characterized by perplexing variability in reported performance, often making results across different language pairs difficult to contextualize. For researchers focused on specific language groups -- such as ancient languages -- it is nearly impossible to determine if breakthroughs reported in other contexts (e.g., native African or American languages) result from superior methodologies or are merely artifacts of benchmark collection. To address this problem, we introduce the FRED Difficulty Metrics, which include the Fertility Ratio (F), Retrieval Proxy (R), Pre-training Exposure (E), and Corpus Diversity (D) and serve as dataset-intrinsic metrics to contextualize reported scores. These metrics reveal that a significant portion of result variability is explained by train-test overlap and pre-training exposure rather than model capability. Additionally, we identify that some languages -- particularly extinct and non-Latin indigenous languages -- suffer from poor tokenization coverage (high token fertility), highlighting a fundamental limitation of transferring models from high-resource languages that lack a shared vocabulary. By providing these indices alongside performance scores, we enable more transparent evaluation of cross-lingual transfer and provide a more reliable foundation for the XLR MT community.

📄 PDF Abstract BibTeX arXiv:2603.25222

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferMachine Translation

Similar Papers 제목 키워드 기반

FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation

2026-04-23 · Jinhee Jang, Juhwan Choi, Dongjin Lee, Seunguk Yu 외 arxiv

Quality Estimation (QE) aims to assess machine translation quality without reference translations, but recent studies have shown that existing QE models exhibit systematic gender bias. In particular, they tend to favor m…

Machine Translation

Calibrating Translation Decoding with Quality Estimation on LLMs

2025-04-26 · Di wu, Yibin Lei, Christof Monz

Neural machine translation (NMT) systems typically employ maximum a posteriori (MAP) decoding to select the highest-scoring translation from the distribution mass. However, recent evidence highlights the inadequacy of MA…

2kMachine TranslationNMTTranslation

Quran-MD: A Fine-Grained Multilingual Multimodal Dataset of the Quran

2026-01-25 · Muhammad Umar Salman, Mohammad Areeb Qazi, Mohammed Talha Alam arxiv

We present Quran MD, a comprehensive multimodal dataset of the Quran that integrates textual, linguistic, and audio dimensions at the verse and word levels. For each verse (ayah), the dataset provides its original Arabic…

Text-To-Speech SynthesisSemantic RetrievalSpeech RecognitionStyle Transfer

Quran Recitation Recognition using End-to-End Deep Learning

2023-05-10 · Ahmad Al Harere, Khloud Al Jallad

The Quran is the holy scripture of Islam, and its recitation is an important aspect of the religion. Recognizing the recitation of the Holy Quran automatically is a challenging task due to its unique rules that are not a…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDeep Learning+2

Rethinking Round-Trip Translation for Machine Translation Evaluation

2022-09-15 · Terry Yue Zhuo, Qiongkai Xu, Xuanli He, Trevor Cohn

Automatic evaluation on low-resource language translation suffers from a deficiency of parallel corpora. Round-trip translation could be served as a clever and straightforward technique to alleviate the requirement of th…

Machine TranslationTranslation