paper-with-me

Papers

Enhancing Human Evaluation in Machine Translation with Comparative Judgment

2025-02-25 · Yixiao Song, Parker Riley, Daniel Deutsch, Markus Freitag

Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. This study explores the integration of comparative judgment into human annotation for machine translation (MT) and evaluates three annotation setups-point-wise Multidimensional Quality Metrics (MQM), side-by-side (SxS) MQM, and its simplified version SxS relative ranking (RR). In MQM, annotators mark error spans with categories and severity levels. SxS MQM extends MQM to pairwise error annotation for two translations of the same input, while SxS RR focuses on selecting the better output without labeling errors. Key findings are: (1) the SxS settings achieve higher inter-annotator agreement than MQM; (2) SxS MQM enhances inter-translation error marking consistency compared to MQM by, on average, 38.5% for explicitly compared MT systems and 19.5% for others; (3) all annotation settings return stable system rankings, with SxS RR offering a more efficient alternative to (SxS) MQM; (4) the SxS settings highlight subtle errors overlooked in MQM without altering absolute system evaluations. To spur further research, we will release the triply annotated datasets comprising 377 ZhEn and 104 EnDe annotation examples.

📄 PDF Abstract BibTeX arXiv:2502.17797

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Does Multimodality Help Human and Machine for Translation and Image Captioning?

2016-05-30 · WS 2016 8 · Ozan Caglayan, Walid Aransa, Yaxing Wang, Marc Masana 외

This paper presents the systems developed by LIUM and CVC for the WMT16 Multimodal Machine Translation challenge. We explored various comparative methods, namely phrase-based systems and attentional recurrent neural netw…

Image CaptioningImage DescriptionMachine TranslationMultimodal Machine Translation+1

A Comparative Study of English-Chinese Translations of Court Texts by Machine and Human Translators and the Word2Vec Based Similarity Measure's Ability To Gauge Human Evaluation Biases

2019-08-01 · WS 2019 8 · Ming Qian, Jessie Liu, Chao-Feng Li, Liming Pals

Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation

2024-01-10 · Zhaokun Jiang, Qianxi Lv, Ziyin Zhang, Lei Lei

Large language models have demonstrated parallel and even superior translation performance compared to neural machine translation (NMT) systems. However, existing comparative studies between them mainly rely on automated…

Machine TranslationNMTPrompt EngineeringTranslation

Machine Translation Quality: A comparative evaluation of SMT, NMT and tailored-NMT outputs

2020-11-01 · EAMT 2020 11 · Maria Stasimioti, Vilelmini Sosoni, Katia Kermanidis, Despoina Mouratidis

The present study aims to compare three systems: a generic statistical machine translation (SMT), a generic neural machine translation (NMT) and a tailored-NMT system focusing on the English to Greek language pair. The c…

Machine TranslationNMTTranslation

Context-Aware Monolingual Human Evaluation of Machine Translation

2025-04-10 · Silvio Picinini, Sheila Castilho

This paper explores the potential of context-aware monolingual human evaluation for assessing machine translation (MT) when no source is given for reference. To this end, we compare monolingual with bilingual evaluations…

Machine TranslationTranslation