paper-with-me

홈 › Papers

Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!

2024-08-25 · Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Edoardo Barba, Roberto Navigli

Annually, at the Conference of Machine Translation (WMT), the Metrics Shared Task organizers conduct the meta-evaluation of Machine Translation (MT) metrics, ranking them according to their correlation with human judgments. Their results guide researchers toward enhancing the next generation of metrics and MT systems. With the recent introduction of neural metrics, the field has witnessed notable advancements. Nevertheless, the inherent opacity of these metrics has posed substantial challenges to the meta-evaluation process. This work highlights two issues with the meta-evaluation framework currently employed in WMT, and assesses their impact on the metrics rankings. To do this, we introduce the concept of sentinel metrics, which are designed explicitly to scrutinize the meta-evaluation process's accuracy, robustness, and fairness. By employing sentinel metrics, we aim to validate our findings, and shed light on and monitor the potential biases or inconsistencies in the rankings. We discover that the present meta-evaluation framework favors two categories of metrics: i) those explicitly trained to mimic human quality assessments, and ii) continuous metrics. Finally, we raise concerns regarding the evaluation capabilities of state-of-the-art metrics, emphasizing that they might be basing their assessments on spurious correlations found in their training data.

📄 PDF Abstract BibTeX arXiv:2408.13831

Code (1)

sapienzanlp/guardians-mt-eval 공식 구현 pytorch

Tasks

FairnessMachine TranslationTranslation

Similar Papers 제목 키워드 기반

Don't Sweat the Small Stuff: Segment-Level Meta-Evaluation Based on Pairwise Difference Correlation

2025-09-29 · Colten DiIanni, Daniel Deutsch arxiv

This paper introduces Pairwise Difference Pearson (PDP), a novel segment-level meta-evaluation metric for Machine Translation (MT) that address limitations in previous Pearson's $ρ$-based and and Kendall's $τ$-based meta…

Machine Translation

Estimating Machine Translation Difficulty

2025-08-13 · Lorenzo Proietti, Stefano Perrella, Vilém Zouhar, Roberto Navigli 외 arxiv

Machine translation quality has steadily improved over the years, achieving near-perfect translations in recent benchmarks. These high-quality outputs make it difficult to distinguish between state-of-the-art models and …

Machine Translation

Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation

2025-09-24 · Behzad Shayegh, Jan-Thorsten Peter, David Vilar, Tobias Domhan 외 arxiv

We investigate the tradeoff between adequacy and fluency in machine translation. We show the severity of this tradeoff at the evaluation level and analyze where popular metrics fall within it. Essentially, current metric…

Machine Translation

MultiEarth 2022 -- The Champion Solution for Image-to-Image Translation Challenge via Generation Models

2022-06-17 · Yuchuan Gou, Bo Peng, Hongchen Liu, Hang Zhou 외

The MultiEarth 2022 Image-to-Image Translation challenge provides a well-constrained test bed for generating the corresponding RGB Sentinel-2 imagery with the given Sentinel-1 VV & VH imagery. In this challenge, we desig…

Image-to-Image TranslationTranslation

MMTE: Corpus and Metrics for Evaluating Machine Translation Quality of Metaphorical Language

2024-06-19 · Shun Wang, Ge Zhang, Han Wu, Tyler Loakman 외

Machine Translation (MT) has developed rapidly since the release of Large Language Models and current MT evaluation is performed through comparison with reference human translations or by predicting quality scores from h…

Machine TranslationTranslation