Evaluating Conjunction Disambiguation on English-to-German and French-to-German WMT 2019 Translation Hypotheses
We present a test set for evaluating an MT system{'}s capability to translate ambiguous conjunctions depending on the sentence structure. We concentrate on the English conjunction {`}but{''} and its French equivalent {}mais{''} which can be translated into two different German conjunctions. We evaluate all English-to-German and French-to-German submissions to the WMT 2019 shared translation task. The evaluation is done mainly automatically, with additional fast manual inspection of unclear cases. All systems almost perfectly recognise the target conjunction {}aber{''}, whereas accuracies for the other target conjunction {}sondern{''} range from 78{\%} to 97{\%}, and the errors are mostly caused by replacing it with the alternative conjunction {}aber{''}. The best performing system for both language pairs is a multilingual Transformer {`}TartuNLP{''} system trained on all WMT 2019 language pairs which use the Latin script, indicating that the multilingual approach is beneficial for conjunction disambiguation. As for other system features, such as using synthetic back-translated data, context-aware, hybrid, etc., no particular (dis)advantages can be observed. Qualitative manual inspection of translation hypotheses shown that highly ranked systems generally produce translations with high adequacy and fluency, meaning that these systems are not only capable of capturing the right conjunction whereas the rest of the translation hypothesis is poor. On the other hand, the low ranked systems generally exhibit lower fluency and poor adequacy.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Sheffield Submissions for WMT18 Multimodal Translation Shared Task
This paper describes the University of Sheffield{'}s submissions to the WMT18 Multimodal Machine Translation shared task. We participated in both tasks 1 and 1b. For task 1, we build on a standard sequence to sequence at…
Data AugmentationMachine TranslationMultimodal Machine TranslationNMT+3Analyzing the Attention Heads for Pronoun Disambiguation in Context-aware Machine Translation Models
In this paper, we investigate the role of attention heads in Context-aware Machine Translation models for pronoun disambiguation in the English-to-German and English-to-French language directions. We analyze their influe…
Machine TranslationSense-Annotated Corpora for Word Sense Disambiguation in Multiple Languages and Domains
The knowledge acquisition bottleneck problem dramatically hampers the creation of sense-annotated data for Word Sense Disambiguation (WSD). Sense-annotated data are scarce for English and almost absent for other language…
Word Sense DisambiguationDISCO: A Large Scale Human Annotated Corpus for Disfluency Correction in Indo-European Languages
Disfluency correction (DC) is the process of removing disfluent elements like fillers, repetitions and corrections from spoken utterances to create readable and interpretable text. DC is a vital post-processing step appl…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+1Automatic Disambiguation of French Discourse Connectives
Discourse connectives (e.g. however, because) are terms that can explicitly convey a discourse relation within a text. While discourse connectives have been shown to be an effective clue to automatically identify discour…