paper-with-me

홈 › Papers

Tilde's Parallel Corpus Filtering Methods for WMT 2018

2018-10-01 · WS 2018 10 · M{\=a}rcis Pinnis

The paper describes parallel corpus filtering methods that allow reducing noise of noisy {``}parallel{''} corpora from a level where the corpora are not usable for neural machine translation training (i.e., the resulting systems fail to achieve reasonable translation quality; well below 10 BLEU points) up to a level where the trained systems show decent (over 20 BLEU points on a 10 million word dataset and up to 30 BLEU points on a 100 million word dataset). The paper also documents Tilde{'}s submissions to the WMT 2018 shared task on parallel corpus filtering.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslationTransliterationWord Alignment

Similar Papers 제목 키워드 기반

Constructing a Chinese---Japanese Parallel Corpus from Wikipedia

2014-05-01 · LREC 2014 5 · Chenhui Chu, Toshiaki Nakazawa, Sadao Kurohashi

Parallel corpora are crucial for statistical machine translation (SMT). However, they are quite scarce for most language pairs, such as Chinese―Japanese. As comparable corpora are far more available, many studies have be…

Machine TranslationSentenceTranslation

Alibaba Submission to the WMT18 Parallel Corpus Filtering Task

2018-10-01 · WS 2018 10 · Jun Lu, Xiaoyu Lv, Yangbin Shi, Boxing Chen

This paper describes the Alibaba Machine Translation Group submissions to the WMT 2018 Shared Task on Parallel Corpus Filtering. While evaluating the quality of the parallel corpus, the three characteristics of the corpu…

DiversityMachine TranslationSentenceTranslation+1

Coverage and Cynicism: The AFRL Submission to the WMT 2018 Parallel Corpus Filtering Task

2018-10-01 · WS 2018 10 · Grant Erdmann, Jeremy Gwinnup

The WMT 2018 Parallel Corpus Filtering Task aims to test various methods of filtering a noisy parallel corpus, to make it useful for training machine translation systems. We describe the AFRL submissions, including their…

Machine TranslationTranslation

Language Richness of the Web

2012-05-01 · LREC 2012 5 · Martin Majli{\v{s}}, Zden{\v{e}}k {\v{Z}}abokrtsk{\'y}

We have built a corpus containing texts in 106 languages from texts available on the Internet and on Wikipedia. The W2C Web Corpus contains 54.7{\textasciitilde}GB of text and the W2C Wiki Corpus contains 8.5{\textasciit…

Alibaba Submission to the WMT20 Parallel Corpus Filtering Task

2020-11-01 · WMT (EMNLP) 2020 11 · Jun Lu, Xin Ge, Yangbin Shi, Yuqi Zhang

This paper describes the Alibaba Machine Translation Group submissions to the WMT 2020 Shared Task on Parallel Corpus Filtering and Alignment. In the filtering task, three main methods are applied to evaluate the quality…

DiversityLanguage IdentificationMachine TranslationSentence+3