paper-with-me

Papers

Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation

2020-05-11 · ACL 2020 6 · Aditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, Yonghui Wu

Over the last few years two promising research directions in low-resource neural machine translation (NMT) have emerged. The first focuses on utilizing high-resource languages to improve the quality of low-resource languages via multilingual NMT. The second direction employs monolingual data with self-supervision to pre-train translation models, followed by fine-tuning on small amounts of supervised data. In this work, we join these two lines of research and demonstrate the efficacy of monolingual data with self-supervision in multilingual NMT. We offer three major results: (i) Using monolingual data significantly boosts the translation quality of low-resource languages in multilingual models. (ii) Self-supervision improves zero-shot translation quality in multilingual models. (iii) Leveraging monolingual data with self-supervision provides a viable path towards adding new languages to multilingual models, getting up to 33 BLEU on ro-en translation without any parallel data or back-translation.

📄 PDF Abstract BibTeX arXiv:2005.04816

Code (0)

등록된 구현이 없습니다.

Tasks

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

2025-09-22 · María Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi, Maureen de Seyssel arxiv

Self-supervised learning (SSL) has made significant advances in speech representation learning. Models like wav2vec 2.0 and HuBERT have achieved state-of-the-art results in tasks such as speech recognition, particularly …

Self-Supervised LearningRepresentation LearningSpeech RecognitionVisual Grounding

Joint Unsupervised and Supervised Training for Multilingual ASR

2021-11-15 · Junwen Bai, Bo Li, Yu Zhang, Ankur Bapna 외

Self-supervised training has shown promising gains in pretraining models and facilitating the downstream finetuning for speech recognition, like multilingual ASR. Most existing methods adopt a 2-stage scheme where the se…

Language ModelingLanguage ModellingMasked Language Modelingspeech-recognition+2

Multilingual JobBERT for Cross-Lingual Job Title Matching

2025-07-29 · Jens-Joris Decorte, Matthias De Lange, Jeroen Van Hautte arxiv

We introduce JobBERT-V3, a contrastive learning-based model for cross-lingual job title matching. Building on the state-of-the-art monolingual JobBERT-V2, our approach extends support to English, German, Spanish, and Chi…

Contrastive Learning

Learning Unsupervised Multilingual Word Embeddings with Incremental Multilingual Hubs

2019-06-01 · NAACL 2019 6 · Geert Heyman, Bregt Verreet, Ivan Vuli{\'c}, Marie-Francine Moens

Recent research has discovered that a shared bilingual word embedding space can be induced by projecting monolingual word embedding spaces from two languages using a self-learning paradigm without any bilingual supervisi…

Bilingual Lexicon InductionCross-Lingual Word EmbeddingsDependency ParsingDocument Classification+3

Learning Multilingual Meta-Embeddings for Code-Switching Named Entity Recognition

2019-08-01 · WS 2019 8 · Genta Indra Winata, Zhaojiang Lin, Pascale Fung

In this paper, we propose Multilingual Meta-Embeddings (MME), an effective method to learn multilingual representations by leveraging monolingual pre-trained embeddings. MME learns to utilize information from these embed…

Language IdentificationMMEnamed-entity-recognitionNamed Entity Recognition+1