paper-with-me

Papers

Consistent Human Evaluation of Machine Translation across Language Pairs

2022-05-17 · AMTA 2022 9 · Daniel Licht, Cynthia Gao, Janice Lam, Francisco Guzman, Mona Diab, Philipp Koehn

Obtaining meaningful quality scores for machine translation systems through human evaluation remains a challenge given the high variability between human evaluators, partly due to subjective expectations for translation quality for different language pairs. We propose a new metric called XSTS that is more focused on semantic equivalence and a cross-lingual calibration method that enables more consistent assessment. We demonstrate the effectiveness of these novel contributions in large scale evaluation studies across up to 14 language pairs, with translation both into and out of English.

📄 PDF Abstract BibTeX arXiv:2205.08533

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Creativity Bias: How Machine Evaluation Struggles with Creativity in Literary Translations

2026-05-13 · Kyo Gerrits, Rik van Noord, Ana Guerberof Arenas arxiv

This article investigates the performance of automatic evaluation metrics (AEMs) and LLM-as-a-judge evaluation on literary translation across multiple languages, genres, and translation modalities. The aim is to assess h…

Machine Translation

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels

2024-11-21 · Jianhao Yan, Pingchuan Yan, Yulong Chen, Jing Li 외

This study presents a comprehensive evaluation of GPT-4's translation capabilities compared to human translators of varying expertise levels. Through systematic human evaluation using the MQM schema, we assess translatio…

BenchmarkingMachine TranslationTranslation

Contrastive ESA: Human Evaluation of Multiple Translations at Once

2026-07-29 · Vilém Zouhar, Roman Grundkiewicz, Sara Rajaee, Parker Riley 외 arxiv

Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and cost. We introduce Contrastive Error Span Annotation (cESA), a protoco…

Machine Translation

How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs

2024-10-24 · Ran Zhang, Wei Zhao, Steffen Eger

Recent research has focused on literary machine translation (MT) as a new challenge in MT. However, the evaluation of literary MT remains an open problem. We contribute to this ongoing discussion by introducing LITEVAL-C…

2kMachine TranslationTranslation

TranslateGemma Technical Report

2026-01-13 · Mara Finkelstein, Isaac Caswell, Tobias Domhan, Jan-Thorsten Peter 외 arxiv

We present TranslateGemma, a suite of open machine translation models based on the Gemma 3 foundation models. To enhance the inherent multilingual capabilities of Gemma 3 for the translation task, we employ a two-stage f…

Reinforcement LearningMachine Translation