paper-with-me

홈 › Papers

Joint Training for Neural Machine Translation Models with Monolingual Data

2018-03-01 · Zhirui Zhang, Shujie Liu, Mu Li, Ming Zhou, Enhong Chen

Monolingual data have been demonstrated to be helpful in improving translation quality of both statistical machine translation (SMT) systems and neural machine translation (NMT) systems, especially in resource-poor or domain adaptation tasks where parallel data are not rich enough. In this paper, we propose a novel approach to better leveraging monolingual data for neural machine translation by jointly learning source-to-target and target-to-source NMT models for a language pair with a joint EM optimization method. The training process starts with two initial NMT models pre-trained on parallel data for each direction, and these two models are iteratively updated by incrementally decreasing translation losses on training data. In each iteration step, both NMT models are first used to translate monolingual data from one language to the other, forming pseudo-training data of the other NMT model. Then two new NMT models are learnt from parallel data together with the pseudo training data. Both NMT models are expected to be improved and better pseudo-training data can be generated in next step. Experiment results on Chinese-English and English-German translation tasks show that our approach can simultaneously improve translation quality of source-to-target and target-to-source models, significantly outperforming strong baseline systems which are enhanced with monolingual data for model training including back-translation.

📄 PDF Abstract BibTeX arXiv:1803.00353

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationMachine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Improving Translation of Out Of Vocabulary Words using Bilingual Lexicon Induction in Low-Resource Machine Translation

2022-09-01 · AMTA 2022 9 · Jonas Waldendorf, Alexandra Birch, Barry Hadow, Antonio Valerio Micele Barone

Dictionary-based data augmentation techniques have been used in the field of domain adaptation to learn words that do not appear in the parallel training data of a machine translation model. These techniques strive to le…

Bilingual Lexicon InductionData AugmentationDomain AdaptationMachine Translation+3

Neural Machine Translation with Monolingual Translation Memory

2021-05-24 · ACL 2021 5 · Deng Cai, Yan Wang, Huayang Li, Wai Lam 외

Prior work has proved that Translation memory (TM) can boost the performance of Neural Machine Translation (NMT). In contrast to existing work that uses bilingual corpus as TM and employs source-side similarity search fo…

Domain AdaptationMachine TranslationNMTRetrieval+1

Multi-task Learning for Multilingual Neural Machine Translation

2020-10-06 · EMNLP 2020 11 · Yiren Wang, ChengXiang Zhai, Hany Hassan Awadalla

While monolingual data has been shown to be useful in improving bilingual neural machine translation (NMT), effectively and efficiently leveraging monolingual data for Multilingual NMT (MNMT) systems is a less explored a…

Cross-Lingual TransferDenoisingMachine TranslationMulti-Task Learning+3

Using Target-side Monolingual Data for Neural Machine Translation through Multi-task Learning

2017-09-01 · EMNLP 2017 9 · Tobias Domhan, Felix Hieber

The performance of Neural Machine Translation (NMT) models relies heavily on the availability of sufficient amounts of parallel data, and an efficient and effective way of leveraging the vastly available amounts of monol…

DecoderLanguage ModelingLanguage ModellingMachine Translation+3

Efficient Unsupervised NMT for Related Languages with Cross-Lingual Language Models and Fidelity Objectives

2021-04-01 · EACL (VarDial) 2021 4 · Rami Aly, Andrew Caines, Paula Buttery

The most successful approach to Neural Machine Translation (NMT) when only monolingual training data is available, called unsupervised machine translation, is based on back-translation where noisy translations are genera…

DenoisingLanguage ModelingLanguage ModellingMachine Translation+3