Improving Low-Resource Neural Machine Translation with Filtered Pseudo-Parallel Corpus
Large-scale parallel corpora are indispensable to train highly accurate machine translators. However, manually constructed large-scale parallel corpora are not freely available in many language pairs. In previous studies, training data have been expanded using a pseudo-parallel corpus obtained using machine translation of the monolingual corpus in the target language. However, in low-resource language pairs in which only low-accuracy machine translation systems can be used, translation quality is reduces when a pseudo-parallel corpus is used naively. To improve machine translation performance with low-resource language pairs, we propose a method to expand the training data effectively via filtering the pseudo-parallel corpus using a quality estimation based on back-translation. As a result of experiments with three language pairs using small, medium, and large parallel corpora, language pairs with fewer training data filtered out more sentence pairs and improved BLEU scores more significantly.
Code (1)
Tasks
Language ModelingLanguage ModellingLow Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationSentenceTranslationSimilar Papers 제목 키워드 기반
Domain Adaptation for NMT via Filtered Iterative Back-Translation
Domain-specific Neural Machine Translation (NMT) model can provide improved performance, however, it is difficult to always access a domain-specific parallel corpus. Iterative Back-Translation can be used for fine-tuning…
Domain AdaptationMachine TranslationNMTTranslationGFST: Gender-Filtered Self-Training for More Accurate Gender in Translation
Targeted evaluations have found that machine translation systems often output incorrect gender in translations, even when the gender is clear from context. Furthermore, these incorrectly gendered translations have the po…
Machine TranslationTranslationAPE-then-QE: Correcting then Filtering Pseudo Parallel Corpora for MT Training Data Creation
Automatic Post-Editing (APE) is the task of automatically identifying and correcting errors in the Machine Translation (MT) outputs. We propose a repair-filter-use methodology that uses an APE system to correct errors on…
Automatic Post-EditingMachine TranslationSentenceTranslation"A Little is Enough": Few-Shot Quality Estimation based Corpus Filtering improves Machine Translation
Quality Estimation (QE) is the task of evaluating the quality of a translation when reference translation is not available. The goal of QE aligns with the task of corpus filtering, where we assign the quality score to th…
Machine TranslationSentenceTransfer LearningTranslationImproving Machine Translation with Phrase Pair Injection and Corpus Filtering
In this paper, we show that the combination of Phrase Pair Injection and Corpus Filtering boosts the performance of Neural Machine Translation (NMT) systems. We extract parallel phrases and sentences from the pseudo-para…
Machine TranslationNMTTranslation