paper-with-me

Papers

Span-Level Machine Translation Meta-Evaluation

2026-03-20 · Stefano Perrella, Eric Morales Agostinho, Hugo Zaragoza arxiv

Machine Translation (MT) and automatic MT evaluation have improved dramatically in recent years, enabling numerous novel applications. Automatic evaluation techniques have evolved from producing scalar quality scores to precisely locating translation errors and assigning them error categories and severity levels. However, it remains unclear how to reliably measure the evaluation capabilities of auto-evaluators that do error detection, as no established technique exists in the literature. This work investigates different implementations of span-level precision, recall, and F-score, showing that seemingly similar approaches can yield substantially different rankings, and that certain widely-used techniques are unsuitable for evaluating MT error detection. We propose "match with partial overlap and partial credit" (MPP) with micro-averaging as a robust meta-evaluation strategy and release code for its use publicly. Finally, we use MPP to assess the state of the art in MT error detection.

📄 PDF Abstract BibTeX arXiv:2603.19921

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation

2025-09-24 · Behzad Shayegh, Jan-Thorsten Peter, David Vilar, Tobias Domhan 외 arxiv

We investigate the tradeoff between adequacy and fluency in machine translation. We show the severity of this tradeoff at the evaluation level and analyze where popular metrics fall within it. Essentially, current metric…

Machine Translation

On Context Span Needed for Machine Translation Evaluation

2020-05-01 · LREC 2020 5 · Sheila Castilho, Maja Popovi{\'c}, Andy Way

Despite increasing efforts to improve evaluation of machine translation (MT) by going beyond the sentence level to the document level, the definition of what exactly constitutes a {``}document level{''} is still not clea…

Machine TranslationSentenceTranslation

xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection

2023-10-16 · Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur 외

Widely used learned metrics for machine translation evaluation, such as COMET and BLEURT, estimate the quality of a translation hypothesis by providing a single sentence-level score. As such, they offer little insight in…

Machine TranslationSentenceTranslation

Design and compilation of a specialized Spanish-German parallel corpus

2012-05-01 · LREC 2012 5 · Carla Parra Escart{\'\i}n

This paper discusses the design and compilation of the TRIS corpus, a specialized parallel corpus of Spanish and German texts. It will be used for phraseological research aimed at improving statistical machine translatio…

Machine TranslationSentenceTranslation

Training and Meta-Evaluating Machine Translation Evaluation Metrics at the Paragraph Level

2023-08-25 · Daniel Deutsch, Juraj Juraska, Mara Finkelstein, Markus Freitag

As research on machine translation moves to translating text beyond the sentence level, it remains unclear how effective automatic evaluation metrics are at scoring longer translations. In this work, we first propose a m…

Machine TranslationSentence