Baidu Neural Machine Translation Systems for WMT19
In this paper we introduce the systems Baidu submitted for the WMT19 shared task on Chinese{\textless}-{\textgreater}English news translation. Our systems are based on the Transformer architecture with some effective improvements. Data selection, back translation, data augmentation, knowledge distillation, domain adaptation, model ensemble and re-ranking are employed and proven effective in our experiments. Our Chinese-{\textgreater}English system achieved the highest case-sensitive BLEU score among all constrained submissions, and our English-{\textgreater}Chinese system ranked the second in all submissions.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationDomain AdaptationKnowledge DistillationMachine TranslationRe-RankingTranslationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Ensemble Sequence Level Training for Multimodal MT: OSU-Baidu WMT18 Multimodal Machine Translation System Report
This paper describes multimodal machine translation systems developed jointly by Oregon State University and Baidu Research for WMT 2018 Shared Task on multimodal translation. In this paper, we introduce a simple approac…
DecoderMachine TranslationMultimodal Machine Translationreinforcement-learning+3Robust Machine Translation with Domain Sensitive Pseudo-Sources: Baidu-OSU WMT19 MT Robustness Shared Task System Report
This paper describes the machine translation system developed jointly by Baidu Research and Oregon State University for WMT 2019 Machine Translation Robustness Shared Task. Translation of social media is a very challengi…
fr-enMachine TranslationTranslationCrafting Adversarial Examples for Neural Machine Translation
Effective adversary generation for neural machine translation (NMT) is a crucial prerequisite for building robust machine translation systems. In this work, we investigate veritable evaluations of NMT adversarial attacks…
Machine TranslationNMTTranslationvalidBSTC: A Large-Scale Chinese-English Speech Translation Dataset
This paper presents BSTC (Baidu Speech Translation Corpus), a large-scale Chinese-English speech translation dataset. This dataset is constructed based on a collection of licensed videos of talks or lectures, including a…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1AlphaMWE: Construction of Multilingual Parallel Corpora with MWE Annotations
In this work, we present the construction of multilingual parallel corpora with annotation of multiword expressions (MWEs). MWEs include verbal MWEs (vMWEs) defined in the PARSEME shared task that have a verb as the head…
Machine TranslationSentenceTranslation