Tilde's Parallel Corpus Filtering Methods for WMT 2018
The paper describes parallel corpus filtering methods that allow reducing noise of noisy {``}parallel{''} corpora from a level where the corpora are not usable for neural machine translation training (i.e., the resulting systems fail to achieve reasonable translation quality; well below 10 BLEU points) up to a level where the trained systems show decent (over 20 BLEU points on a 10 million word dataset and up to 30 BLEU points on a 100 million word dataset). The paper also documents Tilde{'}s submissions to the WMT 2018 shared task on parallel corpus filtering.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationTranslationTransliterationWord AlignmentSimilar Papers 제목 키워드 기반
Constructing a Chinese---Japanese Parallel Corpus from Wikipedia
Parallel corpora are crucial for statistical machine translation (SMT). However, they are quite scarce for most language pairs, such as Chinese―Japanese. As comparable corpora are far more available, many studies have be…
Machine TranslationSentenceTranslationAlibaba Submission to the WMT18 Parallel Corpus Filtering Task
This paper describes the Alibaba Machine Translation Group submissions to the WMT 2018 Shared Task on Parallel Corpus Filtering. While evaluating the quality of the parallel corpus, the three characteristics of the corpu…
DiversityMachine TranslationSentenceTranslation+1Coverage and Cynicism: The AFRL Submission to the WMT 2018 Parallel Corpus Filtering Task
The WMT 2018 Parallel Corpus Filtering Task aims to test various methods of filtering a noisy parallel corpus, to make it useful for training machine translation systems. We describe the AFRL submissions, including their…
Machine TranslationTranslationLanguage Richness of the Web
We have built a corpus containing texts in 106 languages from texts available on the Internet and on Wikipedia. The W2C Web Corpus contains 54.7{\textasciitilde}GB of text and the W2C Wiki Corpus contains 8.5{\textasciit…
Alibaba Submission to the WMT20 Parallel Corpus Filtering Task
This paper describes the Alibaba Machine Translation Group submissions to the WMT 2020 Shared Task on Parallel Corpus Filtering and Alignment. In the filtering task, three main methods are applied to evaluate the quality…
DiversityLanguage IdentificationMachine TranslationSentence+3