paper-with-me

홈 › Papers

Fine-grained linguistic evaluation for state-of-the-art Machine Translation

2020-10-13 · WMT (EMNLP) 2020 11 · Eleftherios Avramidis, Vivien Macketanz, Ursula Strohriegel, Aljoscha Burchardt, Sebastian Möller

This paper describes a test suite submission providing detailed statistics of linguistic performance for the state-of-the-art German-English systems of the Fifth Conference of Machine Translation (WMT20). The analysis covers 107 phenomena organized in 14 categories based on about 5,500 test items, including a manual annotation effort of 45 person hours. Two systems (Tohoku and Huoshan) appear to have significantly better test suite accuracy than the others, although the best system of WMT20 is not significantly better than the one from WMT19 in a macro-average. Additionally, we identify some linguistic phenomena where all systems suffer (such as idioms, resultative predicates and pluperfect), but we are also able to identify particular weaknesses for individual systems (such as quotation marks, lexical ambiguity and sluicing). Most of the systems of WMT19 which submitted new versions this year show improvements.

📄 PDF Abstract BibTeX arXiv:2010.06359

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Linguistic Evaluation for the 2021 State-of-the-art Machine Translation Systems for German to English and English to German

2021-11-01 · WMT (EMNLP) 2021 11 · Vivien Macketanz, Eleftherios Avramidis, Shushen Manakhimova, Sebastian Möller

We are using a semi-automated test suite in order to provide a fine-grained linguistic evaluation for state-of-the-art machine translation systems. The evaluation includes 18 German to English and 18 English to German sy…

Machine TranslationTranslation

Fine-grained evaluation of Quality Estimation for Machine translation based on a linguistically motivated Test Suite

2018-03-01 · WS 2018 3 · Eleftherios Avramidis, Vivien Macketanz, Arle Lommel, Hans Uszkoreit
Automatic Post-EditingCommon Sense ReasoningMachine TranslationTranslation

A Pragmatic Guide to Geoparsing Evaluation

2018-10-29 · Milan Gritta, Mohammad Taher Pilehvar, Nigel Collier

Empirical methods in geoparsing have thus far lacked a standard evaluation framework describing the task, metrics and data used to compare state-of-the-art systems. Evaluation is further made inconsistent, even unreprese…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

Fine-grained evaluation of German-English Machine Translation based on a Test Suite

2019-10-16 · WS 2018 10 · Vivien Macketanz, Eleftherios Avramidis, Aljoscha Burchardt, Hans Uszkoreit

We present an analysis of 16 state-of-the-art MT systems on German-English based on a linguistically-motivated test suite. The test suite has been devised manually by a team of language professionals in order to cover a …

Machine TranslationSentenceTranslation

Linguistic Evaluation of Support Verb Constructions by OpenLogos and Google Translate

2014-05-01 · LREC 2014 5 · Anabela Barreiro, Johanna Monti, Brigitte Orliac, Susanne Preu{\ss} 외

This paper presents a systematic human evaluation of translations of English support verb constructions produced by a rule-based machine translation (RBMT) system (OpenLogos) and a statistical machine translation (SMT) s…

Machine TranslationTranslation