paper-with-me

홈 › Papers

BLEU might be Guilty but References are not Innocent

2020-04-13 · EMNLP 2020 11 · Markus Freitag, David Grangier, Isaac Caswell

The quality of automatic metrics for machine translation has been increasingly called into question, especially for high-quality systems. This paper demonstrates that, while choice of metric is important, the nature of the references is also critical. We study different methods to collect references and compare their value in automated evaluation by reporting correlation with human evaluation for a variety of systems and metrics. Motivated by the finding that typical references exhibit poor diversity, concentrating around translationese language, we develop a paraphrasing task for linguists to perform on existing reference translations, which counteracts this bias. Our method yields higher correlation with human judgment not only for the submissions of WMT 2019 English to German, but also for Back-translation and APE augmented MT output, which have been shown to have low correlation with automatic metrics using standard references. We demonstrate that our methodology improves correlation with all modern evaluation metrics we look at, including embedding-based methods. To complete this picture, we reveal that multi-reference BLEU does not improve the correlation for high quality output, and present an alternative multi-reference formulation that is more effective.

📄 PDF Abstract BibTeX arXiv:2004.06063

Code (2)

google/wmt19-paraphrased-references 공식 구현
atreyasha/semantic-isometry-nmt pytorch

Tasks

DiversityMachine TranslationTranslation

Similar Papers 제목 키워드 기반

Whose Bias?

2021-11-19 · Vasudha Jain, Mark Whitmeyer

Law enforcement acquires costly evidence with the aim of securing the conviction of a defendant, who is convicted if a decision-maker's belief exceeds a certain threshold. Either law enforcement or the decision-maker is …

Hierarchical Quantized Representations for Script Generation

2018-08-28 · EMNLP 2018 10 · Noah Weber, Leena Shekhar, Niranjan Balasubramanian, Nathanael Chambers

Scripts define knowledge about how everyday scenarios (such as going to a restaurant) are expected to unfold. One of the challenges to learning scripts is the hierarchical nature of the knowledge. For example, a suspect …

DecoderLanguage ModelingLanguage ModellingQuantization+1

INNOCENT FAVOUR PRAISE

2025-06-14 · 06/14 2025 6 · INNOCENT FAVOUR PRAISE

SHORT BIOGRAPHY My Name is INNOCENT FAVOUR PRAISE I I'M FROM IMO STATE I WAS BORN 11/10/2009, I was also Born In Ilorin Kwara State, Also Born In The Family Of Mr And Miss INNOCENT, I'M Introduced Into A Business Called…

Marketing

Who Killed Albert Einstein? From Open Data to Murder Mystery Games

2018-02-14 · Gabriella A. B. Barros, Michael Cerny Green, Antonios Liapis, Julian Togelius

This paper presents a framework for generating adventure games from open data. Focusing on the murder mystery type of adventure games, the generator is able to transform open data from Wikipedia articles, OpenStreetMap a…

Articles

Optimizing Voting Order on Sequential Juries: A Median Voter Theorem and Beyond

2020-06-24 · Steve Alpern, Bo Chen

We consider an odd-sized "jury", which votes sequentially between two states of Nature (say A and B, or Innocent and Guilty) with the majority opinion determining the verdict. Jurors have private information in the form …