paper-with-me

홈 › Papers

Error Span Annotation: A Balanced Approach for Human Evaluation of Machine Translation

2024-06-17 · Tom Kocmi, Vilém Zouhar, Eleftherios Avramidis, Roman Grundkiewicz, Marzena Karpinska, Maja Popović, Mrinmaya Sachan, Mariya Shmatova

High-quality Machine Translation (MT) evaluation relies heavily on human judgments. Comprehensive error classification methods, such as Multidimensional Quality Metrics (MQM), are expensive as they are time-consuming and can only be done by experts, whose availability may be limited especially for low-resource languages. On the other hand, just assigning overall scores, like Direct Assessment (DA), is simpler and faster and can be done by translators of any level, but is less reliable. In this paper, we introduce Error Span Annotation (ESA), a human evaluation protocol which combines the continuous rating of DA with the high-level error severity span marking of MQM. We validate ESA by comparing it to MQM and DA for 12 MT systems and one human reference translation (English to German) from WMT23. The results show that ESA offers faster and cheaper annotations than MQM at the same quality level, without the requirement of expensive MQM experts.

📄 PDF Abstract BibTeX arXiv:2406.11580

Code (2)

wmt-conference/ErrorSpanAnnotation 공식 구현
appraisedev/appraise

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

AI-Assisted Human Evaluation of Machine Translation

2024-06-18 · Vilém Zouhar, Tom Kocmi, Mrinmaya Sachan

Annually, research teams spend large amounts of money to evaluate the quality of machine translation systems (WMT, inter alia). This is expensive because it requires a lot of expert human labor. In the recently adopted a…

Machine TranslationTranslation

Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation

2025-12-08 · Boxuan Lyu, Haiyue Song, Hidetaka Kamigaito, Chenchen Ding 외 arxiv

Error Span Detection (ESD) extends automatic machine translation (MT) evaluation by localizing translation errors and labeling their severity. Current generative ESD methods typically use Maximum a Posteriori (MAP) decod…

Machine Translation

Is Human Annotation Necessary? Iterative MBR Distillation for Error Span Detection in Machine Translation

2026-03-13 · Boxuan Lyu, Haiyue Song, Zhi Qu arxiv

Error Span Detection (ESD) is a crucial subtask in Machine Translation (MT) evaluation, aiming to identify the location and severity of translation errors. While fine-tuning models on human-annotated data improves ESD pe…

Machine Translation

Contrastive ESA: Human Evaluation of Multiple Translations at Once

2026-07-29 · Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee, Parker Riley 외 arxiv

Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protoco…

Machine Translation

Enhancing Human Evaluation in Machine Translation with Comparative Judgment

2025-02-25 · Yixiao Song, Parker Riley, Daniel Deutsch, Markus Freitag

Human evaluation is crucial for assessing rapidly evolving language models but is influenced by annotator proficiency and task design. This study explores the integration of comparative judgment into human annotation for…

Machine TranslationTranslation