Attaining the Unattainable? Reassessing Claims of Human Parity in Neural Machine Translation
We reassess a recent study (Hassan et al., 2018) that claimed that machine translation (MT) has reached human parity for the translation of news from Chinese into English, using pairwise ranking and considering three variables that were not taken into account in that previous study: the language in which the source side of the test set was originally written, the translation proficiency of the evaluators, and the provision of inter-sentential context. If we consider only original source text (i.e. not translated from another language, or translationese), then we find evidence showing that human parity has not been achieved. We compare the judgments of professional translators against those of non-experts and discover that those of the experts result in higher inter-annotator agreement and better discrimination between human and machine translations. In addition, we analyse the human translations of the test set and identify important translation issues. Finally, based on these findings, we provide a set of recommendations for future human evaluations of MT.
Code (1)
Tasks
Machine TranslationTranslationSimilar Papers 제목 키워드 기반
Reassessing Claims of Human Parity and Super-Human Performance in Machine Translation at WMT 2019
We reassess the claims of human parity and super-human performance made at the news shared task of WMT 2019 for three translation directions: English-to-German, English-to-Russian and German-to-English. First we identify…
Machine TranslationTranslationOn “Human Parity” and “Super Human Performance” in Machine Translation Evaluation
In this paper, we reassess claims of human parity and super human performance in machine translation. Although these terms have already been discussed, as well as the evaluation protocols used to achieved these conclusio…
Machine TranslationTranslationTranslationese in Machine Translation Evaluation
The term translationese has been used to describe the presence of unusual features of translated text. In this paper, we provide a detailed analysis of the adverse effects of translationese on machine translation evaluat…
Machine TranslationTranslationFact-checking AI-generated news reports: Can LLMs catch their own lies?
In this paper, we evaluate the ability of Large Language Models (LLMs) to assess the veracity of claims in ''news reports'' generated by themselves or other LLMs. Our goal is to determine whether LLMs can effectively fac…
DiagnosticFact CheckingRAGRetrieval-augmented GenerationDialogue Evaluation with Offline Reinforcement Learning
Task-oriented dialogue systems aim to fulfill user goals through natural language interactions. They are ideally evaluated with human users, which however is unattainable to do at every iteration of the development phase…
Dialogue EvaluationOffline RLreinforcement-learningReinforcement Learning+2