paper-with-me

홈 › Papers

Data Contamination in Neural Hieroglyphic Translation: A Reproducibility Study

2026-05-08 · Ammar Toutou, Abdelrahman Harb, Christine Basta arxiv

Ancient and endangered languages pose a unique challenge for NLP: their datasets are inherently scarce, difficult to expand, and built from formulaic corpora -- making data-quality issues especially consequential yet rarely audited. Motivated by the need to understand what current NMT can realistically achieve for such languages, we investigate hieroglyphic-to-German translation, where a recent study reported 61.5 BLEU using fine-tuned M2M-100. Our reproduction yields only 37.0 BLEU with the released model. Investigating this gap, we find 2\% of test targets appear identically in training (16/50; 50\% under 8-gram overlap at 70\% threshold). This contamination inflates scores dramatically: contaminated samples achieve up to 83.8 BLEU / 0.924 COMET-22 versus 30.9--39.2 BLEU / 0.622--0.676 COMET-22 on clean samples across five model configurations spanning two architectures. Document-level decontamination reduces contaminated BLEU by only 4.6 points because 8/16 targets persist via other source documents -- target-level deduplication is required. We release a decontaminated 34-sample test set and establish corrected baselines (30.9--39.2 BLEU), providing a realistic assessment of NMT capability for this endangered writing system.

📄 PDF Abstract BibTeX arXiv:2605.07453

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multi-Task Modeling of Phonographic Languages: Translating Middle Egyptian Hieroglyphs

2019-11-01 · EMNLP (IWSLT) 2019 11 · Philipp Wiesenbach, Stefan Riezler

Machine translation of ancient languages faces a low-resource problem, caused by the limited amount of available textual source data and their translations. We present a multi-task modeling approach to translating Middle…

Machine TranslationMulti-Task LearningPOSPOS Tagging+3

Overestimation in LLM Evaluation: A Controlled Large-Scale Study on Data Contamination's Impact on Machine Translation

2025-01-30 · Muhammed Yusuf Kocyigit, Eleftheria Briakou, Daniel Deutsch, Jiaming Luo 외

Data contamination -- the accidental consumption of evaluation examples within the pre-training data -- can undermine the validity of evaluation benchmarks. In this paper, we present a rigorous analysis of the effects of…

Machine Translation

Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation

2024-02-12 · Federico Ranaldi, Elena Sofia Ruzzetti, Dario Onorati, Leonardo Ranaldi 외

Understanding textual description to generate code seems to be an achieved capability of instruction-following Large Language Models (LLMs) in zero-shot scenario. However, there is a severe possibility that this translat…

Instruction FollowingText to SQLText-To-SQLTranslation

Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora

2026-01-21 · Chaymaa Abbas, Nour Shamaa, Mariette Awad arxiv

Data contamination undermines the validity of Large Language Model evaluation by enabling models to rely on memorized benchmark content rather than true generalization. While prior work has proposed contamination detecti…

When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation

2026-01-28 · David Tan, Pinzhen Chen, Josef van Genabith, Koel Dutta Chowdhury arxiv

Large language models (LLMs) can be benchmark-contaminated, resulting in inflated scores that mask memorization as generalization, and in multilingual settings, this memorization can even transfer to "uncontaminated" lan…

Machine Translation