paper-with-me

홈 › Papers

To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation

2021-07-22 · WMT (EMNLP) 2021 11 · Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, Arul Menezes

Automatic metrics are commonly used as the exclusive tool for declaring the superiority of one machine translation system's quality over another. The community choice of automatic metric guides research directions and industrial developments by deciding which models are deemed better. Evaluating metrics correlations with sets of human judgements has been limited by the size of these sets. In this paper, we corroborate how reliable metrics are in contrast to human judgements on -- to the best of our knowledge -- the largest collection of judgements reported in the literature. Arguably, pairwise rankings of two systems are the most common evaluation tasks in research or deployment scenarios. Taking human judgement as a gold standard, we investigate which metrics have the highest accuracy in predicting translation quality rankings for such system pairs. Furthermore, we evaluate the performance of various metrics across different language pairs and domains. Lastly, we show that the sole use of BLEU impeded the development of improved models leading to bad deployment decisions. We release the collection of 2.3M sentence-level human judgements for 4380 systems for further analysis and replication of our work.

📄 PDF Abstract BibTeX arXiv:2107.10821

Code (2)

MicrosoftTranslator/ToShipOrNotToShip 공식 구현
unbabel/mt-telescope

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

Keep It Private: Unsupervised Privatization of Online Text

2024-05-16 · Calvin Bao, Marine Carpuat

Authorship obfuscation techniques hold the promise of helping people protect their privacy in online communications by automatically rewriting text to hide the identity of the original author. However, obfuscation has be…

Language ModelingLanguage ModellingLarge Language Model

Beyond Surrogates: A Quantitative Analysis for Inter-Metric Relationships

2026-03-08 · Yuanhao Pu, Defu Lian, Enhong Chen arxiv

The Consistency property between surrogate losses and evaluation metrics has been extensively studied to ensure that minimizing a loss leads to metric optimality. However, the direct relationship between different evalua…

Toward Automatic Group Membership Annotation for Group Fairness Evaluation

2024-07-12 · Fumian Chen, Dayu Yang, Hui Fang

With the increasing research attention on fairness in information retrieval systems, more and more fairness-aware algorithms have been proposed to ensure fairness for a sustainable and healthy retrieval ecosystem. Howeve…

FairnessInformation RetrievalRetrieval

T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation

2023-07-12 · NeurIPS 2023 11 · Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li 외

Despite the stunning ability to generate high-quality images by recent text-to-image models, current approaches often struggle to effectively compose objects with different attributes and relationships into a complex and…

AttributeImage GenerationText to Image GenerationText-to-Image Generation

MME-CRS: Multi-Metric Evaluation Based on Correlation Re-Scaling for Evaluating Open-Domain Dialogue

2022-06-19 · Pengfei Zhang, Xiaohui Hu, Kaidong Yu, Jian Wang 외

Automatic open-domain dialogue evaluation is a crucial component of dialogue systems. Recently, learning-based evaluation metrics have achieved state-of-the-art performance in open-domain dialogue evaluation. However, th…

Dialogue EvaluationMME