paper-with-me

Papers

You Can Have Your Data and Balance It Too: Towards Balanced and Efficient Multilingual Models

2022-10-13 · Tomasz Limisiewicz, Dan Malkin, Gabriel Stanovsky

Multilingual models have been widely used for cross-lingual transfer to low-resource languages. However, the performance on these languages is hindered by their underrepresentation in the pretraining data. To alleviate this problem, we propose a novel multilingual training technique based on teacher-student knowledge distillation. In this setting, we utilize monolingual teacher models optimized for their language. We use those teachers along with balanced (sub-sampled) data to distill the teachers' knowledge into a single multilingual student. Our method outperforms standard training methods in low-resource languages and retrains performance on high-resource languages while using the same amount of data. If applied widely, our approach can increase the representation of low-resource languages in NLP systems.

📄 PDF Abstract BibTeX arXiv:2210.07135

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual TransferKnowledge Distillation

Similar Papers 제목 키워드 기반

Prompt Balance Matters: Understanding How Imbalanced Few-Shot Learning Affects Multilingual Sense Disambiguation in LLMs

2025-10-04 · Deshan Sumanathilaka, Nicholas Micallef, Julian Hough arxiv

Recent advances in Large Language Models (LLMs) have significantly reshaped the landscape of Natural Language Processing (NLP). Among the various prompting techniques, few-shot prompting has gained considerable attention…

Word Sense DisambiguationFew-Shot Learning

FuxiTranyu: A Multilingual Large Language Model Trained with Balanced Data

2024-08-12 · Haoran Sun, Renren Jin, Shaoyang Xu, Leiyu Pan 외

Large language models (LLMs) have demonstrated prowess in a wide range of tasks. However, many LLMs exhibit significant performance discrepancies between high- and low-resource languages. To mitigate this challenge, we p…

Language ModelingLanguage ModellingLarge Language Model

GeBioToolkit: Automatic Extraction of Gender-Balanced Multilingual Corpus of Wikipedia Biographies

2019-12-10 · LREC 2020 5 · Marta R. Costa-jussà, Pau Li Lin, Cristina España-Bonet

We introduce GeBioToolkit, a tool for extracting multilingual parallel corpora at sentence level, with document and gender information from Wikipedia biographies. Despite thegender inequalitiespresent in Wikipedia, the t…

Sentence

Tiny Aya: Bridging Scale and Multilingual Depth

2026-03-12 · Alejandro R. Salamanca, Diana Abagyan, Daniel D'souza, Ammar Khairi 외 arxiv

Tiny Aya redefines what a small multilingual language model can achieve. Trained on 70 languages and refined through region-aware posttraining, it delivers state-of-the-art in translation quality, strong multilingual und…

MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models

2026-02-18 · Martin Hyben, Sebastian Kula, Jan Cegin, Jakub Simko 외 arxiv

Large Language Models (LLMs) are beginning to reshape how media professionals verify information, yet automated support for detecting check-worthy claims a key step in the fact-checking process remains limited. We introd…