paper-with-me

Papers

A Set of Recommendations for Assessing Human-Machine Parity in Language Translation

2020-04-03 · Samuel Läubli, Sheila Castilho, Graham Neubig, Rico Sennrich, Qinlan Shen, Antonio Toral

The quality of machine translation has increased remarkably over the past years, to the degree that it was found to be indistinguishable from professional human translation in a number of empirical investigations. We reassess Hassan et al.'s 2018 investigation into Chinese to English news translation, showing that the finding of human-machine parity was owed to weaknesses in the evaluation design - which is currently considered best practice in the field. We show that the professional human translations contained significantly fewer errors, and that perceived quality in human evaluation depends on the choice of raters, the availability of linguistic context, and the creation of reference translations. Our results call for revisiting current best practices to assess strong machine translation systems in general and human-machine parity in particular, for which we offer a set of recommendations based on our empirical findings.

📄 PDF Abstract BibTeX arXiv:2004.01694

Code (1)

ZurichNLP/mt-parity-assessment-data 공식 구현

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Reassessing Claims of Human Parity and Super-Human Performance in Machine Translation at WMT 2019

2020-05-12 · EAMT 2020 11 · Antonio Toral

We reassess the claims of human parity and super-human performance made at the news shared task of WMT 2019 for three translation directions: English-to-German, English-to-Russian and German-to-English. First we identify…

Machine TranslationTranslation

Attaining the Unattainable? Reassessing Claims of Human Parity in Neural Machine Translation

2018-08-30 · WS 2018 10 · Antonio Toral, Sheila Castilho, Ke Hu, Andy Way

We reassess a recent study (Hassan et al., 2018) that claimed that machine translation (MT) has reached human parity for the translation of news from Chinese into English, using pairwise ranking and considering three var…

Machine TranslationTranslation

Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation

2018-08-21 · EMNLP 2018 10 · Samuel Läubli, Rico Sennrich, Martin Volk

Recent research suggests that neural machine translation achieves parity with professional human translation on the WMT Chinese--English news translation task. We empirically test this claim with alternative evaluation p…

Machine TranslationSentenceTranslation

Recommendation or Discrimination?: Quantifying Distribution Parity in Information Retrieval Systems

2019-09-13 · Rinat Khaziev, Bryce Casavant, Pearce Washabaugh, Amy A. Winecoff 외

Information retrieval (IR) systems often leverage query data to suggest relevant items to users. This introduces the possibility of unfairness if the query (i.e., input) and the resulting recommendations unintentionally …

Information RetrievalRetrieval

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?

2026-05-11 · Maximilian Triebel, Marco Menner, Dominik Helfenstein arxiv

Vision-Language(-Action) Models (VLMs) are increasingly applied to interactive environments, yet existing benchmarks often overlook the complex physical reasoning required for point-and-click puzzle games. This paper int…

Logical ReasoningVisual Grounding