paper-with-me

홈 › Papers

Lifting the Curse of Multilinguality by Pre-training Modular Transformers

2022-05-12 · NAACL 2022 7 · Jonas Pfeiffer, Naman Goyal, Xi Victoria Lin, Xian Li, James Cross, Sebastian Riedel, Mikel Artetxe

Multilingual pre-trained models are known to suffer from the curse of multilinguality, which causes per-language performance to drop as they cover more languages. We address this issue by introducing language-specific modules, which allows us to grow the total capacity of the model, while keeping the total number of trainable parameters per language constant. In contrast with prior work that learns language-specific components post-hoc, we pre-train the modules of our Cross-lingual Modular (X-Mod) models from the start. Our experiments on natural language inference, named entity recognition and question answering show that our approach not only mitigates the negative interference between languages, but also enables positive transfer, resulting in improved monolingual and cross-lingual performance. Furthermore, our approach enables adding languages post-hoc with no measurable drop in performance, no longer limiting the model usage to the set of pre-trained languages.

📄 PDF Abstract BibTeX arXiv:2205.06266

Code (0)

등록된 구현이 없습니다.

Tasks

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language InferenceQuestion Answering

Similar Papers 제목 키워드 기반

Lifting the Curse of Multilinguality by Pre-training Modular Transformers

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Multilingual pre-trained models are known to suffer from the curse of multilinguality, which causes per-language performance to drop as they cover more languages. We address this issue by introducing language-specific mo…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language Inference+1

No Train but Gain: Language Arithmetic for training-free Language Adapters enhancement

2024-04-24 · Mateusz Klimaszewski, Piotr Andruszkiewicz, Alexandra Birch

Modular deep learning is the state-of-the-art solution for lifting the curse of multilinguality, preventing the impact of negative interference and enabling cross-lingual performance in Multilingual Pre-trained Language …

Task ArithmeticTransfer Learning

Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual Alignment

2024-07-20 · Yongxin Huang, Kexin Wang, Goran Glavaš, Iryna Gurevych

Multilingual sentence encoders are commonly obtained by training multilingual language models to map sentences from different languages into a shared semantic space. As such, they are subject to curse of multilinguality,…

Contrastive LearningMultiple-choiceSentenceSentence Embeddings+1

Multilingual Large Language Models and Curse of Multilinguality

2024-06-15 · Daniil Gurgurov, Tanja Bäumel, Tatiana Anikina

Multilingual Large Language Models (LLMs) have gained large popularity among Natural Language Processing (NLP) researchers and practitioners. These models, trained on huge datasets, show proficiency across various langua…

DecoderXLM-R

When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages

2023-11-15 · Tyler A. Chang, Catherine Arnett, Zhuowen Tu, Benjamin K. Bergen

Multilingual language models are widely used to extend NLP systems to low-resource languages. However, concrete evidence for the effects of multilinguality on language modeling performance in individual languages remains…

Language ModelingLanguage Modelling