Robust Machine Translation with Domain Sensitive Pseudo-Sources: Baidu-OSU WMT19 MT Robustness Shared Task System Report
This paper describes the machine translation system developed jointly by Baidu Research and Oregon State University for WMT 2019 Machine Translation Robustness Shared Task. Translation of social media is a very challenging problem, since its style is very different from normal parallel corpora (e.g. News) and also include various types of noises. To make it worse, the amount of social media parallel corpora is extremely limited. In this paper, we use a domain sensitive training method which leverages a large amount of parallel data from popular domains together with a little amount of parallel data from social media. Furthermore, we generate a parallel dataset with pseudo noisy source sentences which are back-translated from monolingual data using a model trained by a similar domain sensitive way. We achieve more than 10 BLEU improvement in both En-Fr and Fr-En translation compared with the baseline methods.
Code (0)
등록된 구현이 없습니다.
Tasks
fr-enMachine TranslationTranslationSimilar Papers 제목 키워드 기반
Domain Adaptation of Neural Machine Translation by Lexicon Induction
It has been previously noted that neural machine translation (NMT) is very sensitive to domain shift. In this paper, we argue that this is a dual effect of the highly lexicalized nature of NMT, resulting in failure for s…
Domain AdaptationMachine TranslationNMTTranslationThe FISKM\"O Project: Resources and Tools for Finnish-Swedish Machine Translation and Cross-Linguistic Research
This paper presents FISKM{\"O}, a project that focuses on the development of resources and tools for cross-linguistic research and machine translation between Finnish and Swedish. The goal of the project is the compilati…
Machine TranslationTranslationPseudo-Label Training and Model Inertia in Neural Machine Translation
Like many other machine learning applications, neural machine translation (NMT) benefits from over-parameterized deep neural models. However, these models have been observed to be brittle: NMT model predictions are sensi…
Knowledge DistillationMachine TranslationNMTPseudo Label+1Joint Speech Transcription and Translation: Pseudo-Labeling with Out-of-Distribution Data
Self-training has been shown to be helpful in addressing data scarcity for many domains, including vision, speech, and language. Specifically, self-training, or pseudo-labeling, labels unsupervised data and adds that to …
Data AugmentationPseudo LabelPseudo Label FilteringTranslationBoosting Unsupervised Machine Translation with Pseudo-Parallel Data
Even with the latest developments in deep learning and large-scale language modeling, the task of machine translation (MT) of low-resource languages remains a challenge. Neural MT systems can be trained in an unsupervise…
Language ModelingLanguage ModellingMachine TranslationSentence+2