paper-with-me

Papers

Exploiting Out-of-Domain Parallel Data through Multilingual Transfer Learning for Low-Resource Neural Machine Translation

2019-07-06 · WS 2019 8 · Aizhan Imankulova, Raj Dabre, Atsushi Fujita, Kenji Imamura

This paper proposes a novel multilingual multistage fine-tuning approach for low-resource neural machine translation (NMT), taking a challenging Japanese--Russian pair for benchmarking. Although there are many solutions for low-resource scenarios, such as multilingual NMT and back-translation, we have empirically confirmed their limited success when restricted to in-domain data. We therefore propose to exploit out-of-domain data through transfer learning, by using it to first train a multilingual NMT model followed by multistage fine-tuning on in-domain parallel and back-translated pseudo-parallel data. Our approach, which combines domain adaptation, multilingualism, and back-translation, helps improve the translation quality by more than 3.7 BLEU points, over a strong baseline, for this extremely low-resource scenario.

📄 PDF Abstract BibTeX arXiv:1907.03060

Code (1)

aizhanti/jarunc

Tasks

BenchmarkingDomain AdaptationLow Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMTTransfer LearningTranslation

Similar Papers 제목 키워드 기반

Exploiting Domain-Specific Parallel Data on Multilingual Language Models for Low-resource Language Translation

2024-12-27 · Surangika Ranathungaa, Shravan Nayak, Shih-Ting Cindy Huang, Yanke Mao 외

Neural Machine Translation (NMT) systems built on multilingual sequence-to-sequence Language Models (msLMs) fail to deliver expected results when the amount of parallel data for a language, as well as the language's repr…

Machine TranslationNMT

A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models

2024-06-29 · Peiqin Lin, André F. T. Martins, Hinrich Schütze

Recent studies have highlighted the potential of exploiting parallel corpora to enhance multilingual large language models, improving performance in both bilingual tasks, e.g., machine translation, and general-purpose ta…

Language IdentificationMachine TranslationSentencetext-classification+2

Exploiting Multilingualism through Multistage Fine-Tuning for Low-Resource Neural Machine Translation

2019-11-01 · IJCNLP 2019 11 · Raj Dabre, Atsushi Fujita, Chenhui Chu

This paper highlights the impressive utility of multi-parallel corpora for transfer learning in a one-to-many low-resource neural machine translation (NMT) setting. We report on a systematic comparison of multistage fine…

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMT+2

PARADISE: Exploiting Parallel Data for Multilingual Sequence-to-Sequence Pretraining

2021-08-04 · NAACL 2022 7 · Machel Reid, Mikel Artetxe

Despite the success of multilingual sequence-to-sequence pretraining, most existing approaches rely on monolingual corpora, and do not make use of the strong cross-lingual signal contained in parallel data. In this paper…

Cross-Lingual Natural Language InferenceDenoisingMachine TranslationNatural Language Inference+1

PARADISE”:" Exploiting Parallel Data for Multilingual Sequence-to-Sequence Pretraining

2022-05-01 · RepL4NLP (ACL) 2022 5 · Machel Reid, Mikel Artetxe

Despite the success of multilingual sequence-to-sequence pretraining, most existing approaches rely on monolingual corpora and do not make use of the strong cross-lingual signal contained in parallel data. In this paper,…

Cross-Lingual Natural Language InferenceDenoisingMachine TranslationNatural Language Inference+1