paper-with-me

Papers

Exploiting Multilingualism through Multistage Fine-Tuning for Low-Resource Neural Machine Translation

2019-11-01 · IJCNLP 2019 11 · Raj Dabre, Atsushi Fujita, Chenhui Chu

This paper highlights the impressive utility of multi-parallel corpora for transfer learning in a one-to-many low-resource neural machine translation (NMT) setting. We report on a systematic comparison of multistage fine-tuning configurations, consisting of (1) pre-training on an external large (209k{--}440k) parallel corpus for English and a helping target language, (2) mixed pre-training or fine-tuning on a mixture of the external and low-resource (18k) target parallel corpora, and (3) pure fine-tuning on the target parallel corpora. Our experiments confirm that multi-parallel corpora are extremely useful despite their scarcity and content-wise redundancy thus exhibiting the true power of multilingualism. Even when the helping target language is not one of the target languages of our concern, our multistage fine-tuning can give 3{--}9 BLEU score gains over a simple one-to-one model.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMTTransfer LearningTranslation

Similar Papers 제목 키워드 기반

Exploiting Out-of-Domain Parallel Data through Multilingual Transfer Learning for Low-Resource Neural Machine Translation

2019-07-06 · WS 2019 8 · Aizhan Imankulova, Raj Dabre, Atsushi Fujita, Kenji Imamura

This paper proposes a novel multilingual multistage fine-tuning approach for low-resource neural machine translation (NMT), taking a challenging Japanese--Russian pair for benchmarking. Although there are many solutions …

BenchmarkingDomain AdaptationLow Resource Neural Machine TranslationLow-Resource Neural Machine Translation+4

NICT's participation to WAT 2019: Multilingualism and Multi-step Fine-Tuning for Low Resource NMT

2019-11-01 · WS 2019 11 · Raj Dabre, Eiichiro Sumita

In this paper we describe our submissions to WAT 2019 for the following tasks: English{--}Tamil translation and Russian{--}Japanese translation. Our team,{``}NICT-5{''}, focused on multilingual domain adaptation and back…

Domain AdaptationLow Resource NMTNMTTranslation

Multistage Fine-tuning Strategies for Automatic Speech Recognition in Low-resource Languages

2024-11-07 · Leena G Pillai, Kavya Manohar, Basil K Raju, Elizabeth Sherly

This paper presents a novel multistage fine-tuning strategy designed to enhance automatic speech recognition (ASR) performance in low-resource languages using OpenAI's Whisper model. In this approach we aim to build ASR …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

How Vocabulary Sharing Facilitates Multilingualism in LLaMA?

2023-11-15 · Fei Yuan, Shuai Yuan, Zhiyong Wu, Lei LI

Large Language Models (LLMs), often show strong performance on English tasks, while exhibiting limitations on other languages. What is an LLM's multilingual capability when it is trained only on certain languages? The un…

Coursera Corpus Mining and Multistage Fine-Tuning for Improving Lectures Translation

2019-12-26 · LREC 2020 5 · Haiyue Song, Raj Dabre, Atsushi Fujita, Sadao Kurohashi

Lectures translation is a case of spoken language translation and there is a lack of publicly available parallel corpora for this purpose. To address this, we examine a language independent framework for parallel corpus …

BenchmarkingDomain AdaptationMachine TranslationParallel Corpus Mining+2