paper-with-me

홈 › Papers

MetaMetrics-MT: Tuning Meta-Metrics for Machine Translation via Human Preference Calibration

2024-11-01 · David Anugraha, Garry Kuwanto, Lucky Susanto, Derry Tanti Wijaya, Genta Indra Winata

We present MetaMetrics-MT, an innovative metric designed to evaluate machine translation (MT) tasks by aligning closely with human preferences through Bayesian optimization with Gaussian Processes. MetaMetrics-MT enhances existing MT metrics by optimizing their correlation with human judgments. Our experiments on the WMT24 metric shared task dataset demonstrate that MetaMetrics-MT outperforms all existing baselines, setting a new benchmark for state-of-the-art performance in the reference-based setting. Furthermore, it achieves comparable results to leading metrics in the reference-free setting, offering greater efficiency.

📄 PDF Abstract BibTeX arXiv:2411.00390

Code (1)

meta-metrics/metametrics 공식 구현

Tasks

Bayesian OptimizationGaussian ProcessesMachine TranslationTranslation

Similar Papers 제목 키워드 기반

MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

2024-10-03 · Genta Indra Winata, David Anugraha, Lucky Susanto, Garry Kuwanto 외

Understanding the quality of a performance evaluation metric is crucial for ensuring that model outputs align with human preferences. However, it remains unclear how well each metric captures the diverse aspects of these…

Guardians of the Machine Translation Meta-Evaluation: Sentinel Metrics Fall In!

2024-08-25 · Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Edoardo Barba 외

Annually, at the Conference of Machine Translation (WMT), the Metrics Shared Task organizers conduct the meta-evaluation of Machine Translation (MT) metrics, ranking them according to their correlation with human judgmen…

FairnessMachine TranslationTranslation

Feeding Two Birds or Favoring One? Adequacy-Fluency Tradeoffs in Evaluation and Meta-Evaluation of Machine Translation

2025-09-24 · Behzad Shayegh, Jan-Thorsten Peter, David Vilar, Tobias Domhan 외 arxiv

We investigate the tradeoff between adequacy and fluency in machine translation. We show the severity of this tradeoff at the evaluation level and analyze where popular metrics fall within it. Essentially, current metric…

Machine Translation

A New Family of Near-metrics for Universal Similarity

2017-07-21 · Chu Wang, Iraj Saniee, William S. Kennedy, Chris A. White

We propose a family of near-metrics based on local graph diffusion to capture similarity for a wide class of data sets. These quasi-metametrics, as their names suggest, dispense with one or two standard axioms of metric …

Training and Meta-Evaluating Machine Translation Evaluation Metrics at the Paragraph Level

2023-08-25 · Daniel Deutsch, Juraj Juraska, Mara Finkelstein, Markus Freitag

As research on machine translation moves to translating text beyond the sentence level, it remains unclear how effective automatic evaluation metrics are at scoring longer translations. In this work, we first propose a m…

Machine TranslationSentence