paper-with-me

Papers

Taking MT Evaluation Metrics to Extremes: Beyond Correlation with Human Judgments

2019-09-01 · CL 2019 9 · Marina Fomicheva, Lucia Specia

Automatic Machine Translation (MT) evaluation is an active field of research, with a handful of new metrics devised every year. Evaluation metrics are generally benchmarked against manual assessment of translation quality, with performance measured in terms of overall correlation with human scores. Much work has been dedicated to the improvement of evaluation metrics to achieve a higher correlation with human judgments. However, little insight has been provided regarding the weaknesses and strengths of existing approaches and their behavior in different settings. In this work we conduct a broad meta-evaluation study of the performance of a wide range of evaluation metrics focusing on three major aspects. First, we analyze the performance of the metrics when faced with different levels of translation quality, proposing a local dependency measure as an alternative to the standard, global correlation coefficient. We show that metric performance varies significantly across different levels of MT quality: Metrics perform poorly when faced with low-quality translations and are not able to capture nuanced quality distinctions. Interestingly, we show that evaluating low-quality translations is also more challenging for humans. Second, we show that metrics are more reliable when evaluating neural MT than the traditional statistical MT systems. Finally, we show that the difference in the evaluation accuracy for different metrics is maintained even if the gold standard scores are based on different criteria.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Layer or Representation Space: What makes BERT-based Evaluation Metrics Robust?

2022-09-06 · COLING 2022 10 · Doan Nam Long Vu, Nafise Sadat Moosavi, Steffen Eger

The evaluation of recent embedding-based evaluation metrics for text generation is primarily based on measuring their correlation with human evaluations on standard benchmarks. However, these benchmarks are mostly from s…

Text GenerationWord Embeddings

Capturing Unseen Spatial Heat Extremes Through Dependence-Aware Generative Modeling

2025-07-12 · Xinyue Liu, Xiao Peng, Shuyue Yan, Yuntian Chen 외 arxiv

Observed records of climate extremes provide an incomplete view of risk, missing "unseen" events beyond historical experience. Ignoring spatial dependence further underestimates hazards striking multiple locations simult…

Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

2025-03-06 · Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang 외

Spoken dialogue modeling poses challenges beyond text-based language modeling, requiring real-time interaction, turn-taking, and backchanneling. While most Spoken Dialogue Models (SDMs) operate in half-duplex mode-proces…

Language ModelingLanguage ModellingManagement

Kunyu: A High-Performing Global Weather Model Beyond Regression Losses

2023-12-04 · Zekun Ni

Over the past year, data-driven global weather forecasting has emerged as a new alternative to traditional numerical weather prediction. This innovative approach yields forecasts of comparable accuracy at a tiny fraction…

regressionWeather Forecasting

Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics

2024-10-07 · Stefano Perrella, Lorenzo Proietti, Pere-Lluís Huguet Cabot, Edoardo Barba 외

Machine Translation (MT) evaluation metrics assess translation quality automatically. Recently, researchers have employed MT metrics for various new use cases, such as data filtering and translation re-ranking. However, …

Machine TranslationRe-RankingTranslation