paper-with-me

홈 › Papers

Identifying Weaknesses in Machine Translation Metrics Through Minimum Bayes Risk Decoding: A Case Study for COMET

2022-02-10 · Chantal Amrhein, Rico Sennrich

Neural metrics have achieved impressive correlation with human judgements in the evaluation of machine translation systems, but before we can safely optimise towards such metrics, we should be aware of (and ideally eliminate) biases toward bad translations that receive high scores. Our experiments show that sample-based Minimum Bayes Risk decoding can be used to explore and quantify such weaknesses. When applying this strategy to COMET for en-de and de-en, we find that COMET models are not sensitive enough to discrepancies in numbers and named entities. We further show that these biases are hard to fully remove by simply training on additional synthetic data and release our code and data for facilitating further experiments.

📄 PDF Abstract BibTeX arXiv:2202.05148

Code (1)

zurichnlp/mbr-sensitivity 공식 구현

Tasks

de-enMachine TranslationTranslation

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Breeding Machine Translations: Evolutionary approach to survive and thrive in the world of automated evaluation

2023-05-30 · Josef Jon, Ondřej Bojar

We propose a genetic algorithm (GA) based method for modifying n-best lists produced by a machine translation (MT) system. Our method offers an innovative approach to improving MT quality and identifying weaknesses in ev…

Machine TranslationTranslation

Generating Difficult-to-Translate Texts

2025-09-30 · Vilém Zouhar, Wenda Xu, Parker Riley, Juraj Juraska 외 arxiv

Machine translation benchmarks sourced from the real world are quickly obsoleted, due to most examples being easy for state-of-the-art translation models. This limits the benchmark's ability to distinguish which model is…

Machine Translation

How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation

2016-03-25 · EMNLP 2016 11 · Chia-Wei Liu, Ryan Lowe, Iulian V. Serban, Michael Noseworthy 외

We investigate evaluation metrics for dialogue response generation systems where supervised labels, such as task completion, are not available. Recent works in response generation have adopted metrics from machine transl…

Machine TranslationResponse GenerationTranslation

Ties Matter: Meta-Evaluating Modern Metrics with Pairwise Accuracy and Tie Calibration

2023-05-23 · Daniel Deutsch, George Foster, Markus Freitag

Kendall's tau is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations. Its focus on pairwise score comparisons is intuitive but raises the question of how ti…

Machine Translation

Comparison of SMT and RBMT; The Requirement of Hybridization for Marathi-Hindi MT

2017-03-10 · Sreelekha. S, Pushpak Bhattacharyya

We present in this paper our work on comparison between Statistical Machine Translation (SMT) and Rule-based machine translation for translation from Marathi to Hindi. Rule Based systems although robust take lots of time…

Machine TranslationTranslation