paper-with-me

Papers

Investigating Backtranslation in Neural Machine Translation

2018-04-17 · Alberto Poncelas, Dimitar Shterionov, Andy Way, Gideon Maillette de Buy Wenniger, Peyman Passban

A prerequisite for training corpus-based machine translation (MT) systems -- either Statistical MT (SMT) or Neural MT (NMT) -- is the availability of high-quality parallel data. This is arguably more important today than ever before, as NMT has been shown in many studies to outperform SMT, but mostly when large parallel corpora are available; in cases where data is limited, SMT can still outperform NMT. Recently researchers have shown that back-translating monolingual data can be used to create synthetic parallel corpora, which in turn can be used in combination with authentic parallel data to train a high-quality NMT system. Given that large collections of new parallel text become available only quite rarely, backtranslation has become the norm when building state-of-the-art NMT systems, especially in resource-poor scenarios. However, we assert that there are many unknown factors regarding the actual effects of back-translated data on the translation capabilities of an NMT model. Accordingly, in this work we investigate how using back-translated data as a training corpus -- both as a separate standalone dataset as well as combined with human-generated parallel data -- affects the performance of an NMT model. We use incrementally larger amounts of back-translated data to train a range of NMT systems for German-to-English, and analyse the resulting translation performance.

📄 PDF Abstract BibTeX arXiv:1804.06189

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Boosting Neural Machine Translation from Finnish to Northern Sámi with Rule-Based Backtranslation

2021-05-01 · NoDaLiDa 2021 5 · Mikko Aulamo, Sami Virpioja, Yves Scherrer, Jörg Tiedemann

We consider a low-resource translation task from Finnish into Northern Sámi. Collecting all available parallel data between the languages, we obtain around 30,000 sentence pairs. However, there exists a significantly lar…

Machine TranslationNMTSentenceTranslation

Backtranslation in Neural Morphological Inflection

2021-11-01 · EMNLP (insights) 2021 11 · Ling Liu, Mans Hulden

Backtranslation is a common technique for leveraging unlabeled data in low-resource scenarios in machine translation. The method is directly applicable to morphological inflection generation if unlabeled word forms are a…

Machine TranslationMorphological InflectionTranslation

Exploring Parameter-Efficient Fine-Tuning and Backtranslation for the WMT 25 General Translation Task

2025-11-15 · Felipe Fujita, Hideyuki Takada arxiv

In this paper, we explore the effectiveness of combining fine-tuning and backtranslation on a small Japanese corpus for neural machine translation. Starting from a baseline English{\textrightarrow}Japanese model (COMET =…

parameter-efficient fine-tuningMachine Translation

Study on Unsupervised Statistical Machine Translation for Backtranslation

2019-09-01 · RANLP 2019 9 · Anush Kumar, Nihal V. Nayak, Ch, Aditya ra 외

Machine Translation systems have drastically improved over the years for several language pairs. Monolingual data is often used to generate synthetic sentences to augment the training data which has shown to improve the …

Machine TranslationTranslationUnsupervised Machine Translation

Leveraging backtranslation to improve machine translation for Gaelic languages

2019-08-01 · WS 2019 8 · Meghan Dowling, Teresa Lynn, Andy Way
Machine TranslationTranslation