paper-with-me

홈 › Papers

Filtering of Noisy Parallel Corpora Based on Hypothesis Generation

2019-08-01 · WS 2019 8 · Zuzanna Parcheta, Germ{\'a}n Sanchis-Trilles, Francisco Casacuberta

The filtering task of noisy parallel corpora in WMT2019 aims to challenge participants to create filtering methods to be useful for training machine translation systems. In this work, we introduce a noisy parallel corpora filtering system based on generating hypotheses by means of a translation model. We train translation models in both language pairs: Nepali{--}English and Sinhala{--}English using provided parallel corpora. We select the training subset for three language pairs (Nepali, Sinhala and Hindi to English) jointly using bilingual cross-entropy selection to create the best possible translation model for both language pairs. Once the translation models are trained, we translate the noisy corpora and generate a hypothesis for each sentence pair. We compute the smoothed BLEU score between the target sentence and generated hypothesis. In addition, we apply several rules to discard very noisy or inadequate sentences which can lower the translation score. These heuristics are based on sentence length, source and target similarity and source language detection. We compare our results with the baseline published on the shared task website, which uses the Zipporah model, over which we achieve significant improvements in one of the conditions in the shared task. The designed filtering system is domain independent and all experiments are conducted using neural machine translation.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationSentenceTranslation

Similar Papers 제목 키워드 기반

Quality and Coverage: The AFRL Submission to the WMT19 Parallel Corpus Filtering for Low-Resource Conditions Task

2019-08-01 · WS 2019 8 · Grant Erdmann, Jeremy Gwinnup

The WMT19 Parallel Corpus Filtering For Low-Resource Conditions Task aims to test various methods of filtering a noisy parallel corpora, to make them useful for training machine translation systems. This year the noisy c…

Machine TranslationTranslation

Dual Conditional Cross-Entropy Filtering of Noisy Parallel Corpora

2018-09-01 · EMNLP 2018 10 · Marcin Junczys-Dowmunt

In this work we introduce dual conditional cross-entropy filtering for noisy parallel data. For each sentence pair of the noisy parallel corpus we compute cross-entropy scores according to two inverse translation models …

SentenceTranslation

Tilde's Parallel Corpus Filtering Methods for WMT 2018

2018-10-01 · WS 2018 10 · M{\=a}rcis Pinnis

The paper describes parallel corpus filtering methods that allow reducing noise of noisy {``}parallel{''} corpora from a level where the corpora are not usable for neural machine translation training (i.e., the resulting…

Machine TranslationTranslationTransliterationWord Alignment

Accurate semantic textual similarity for cleaning noisy parallel corpora using semantic machine translation evaluation metric: The NRC supervised submissions to the Parallel Corpus Filtering task

2018-10-01 · WS 2018 10 · Chi-kiu Lo, Michel Simard, Darlene Stewart, Samuel Larkin 외

We present our semantic textual similarity approach in filtering a noisy web crawled parallel corpus using YiSi{---}a novel semantic machine translation evaluation metric. The systems mainly based on this supervised appr…

Machine TranslationSemantic Textual SimilarityTranslation

An Unsupervised System for Parallel Corpus Filtering

2018-10-01 · WS 2018 10 · Viktor Hangya, Alex Fraser, er

In this paper we describe LMU Munich{'}s submission for the \textit{WMT 2018 Parallel Corpus Filtering} shared task which addresses the problem of cleaning noisy parallel corpora. The task of mining and cleaning parallel…

Domain AdaptationLanguage ModelingLanguage ModellingMachine Translation+5