paper-with-me

홈 › Papers

XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs

2025-12-11 · Iñaki Lacunza, José Javier Saiz, Alexander Shvets, Aitor Gonzalez-Agirre, Marta Villegas arxiv

Current large language models (LLMs) are trained on massive amounts of text data, primarily from a few dominant languages. Studies suggest that this over-reliance on high-resource languages, such as English, hampers LLM performance in mid- and low-resource languages. To mitigate this problem, we propose to (i) optimize the language distribution by training a small proxy model within a domain-reweighing DoGE algorithm that we extend to XDoGE for a multilingual setup, and (ii) rescale the data and train a full-size model with the established language weights either from scratch or within a continual pre-training phase (CPT). We target six languages possessing a variety of geographic and intra- and inter-language-family relations, namely, English and Spanish (high-resource), Portuguese and Catalan (mid-resource), Galician and Basque (low-resource). We experiment with Salamandra-2b, which is a promising model for these languages. We investigate the effects of substantial data repetition on minor languages and under-sampling on dominant languages using the IberoBench framework for quantitative evaluation. Finally, we release a new promising IberianLLM-7B-Instruct model centering on Iberian languages and English that we pretrained from scratch and further improved using CPT with the XDoGE weights.

📄 PDF Abstract BibTeX arXiv:2512.10545

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models

2025-11-10 · Pukang Ye, Junwei Luo, Xiaolei Dong, Yunbo Yang arxiv

Data duplication within large-scale corpora often impedes large language models' (LLMs) performance and privacy. In privacy-concerned federated learning scenarios, conventional deduplication methods typically rely on tru…

Federated Learning

Logit Reweighting for Topic-Focused Summarization

2025-07-07 · Joschka Braun, Bálint Mucsányi, Seyed Ali Bahrainian arxiv

Generating abstractive summaries that adhere to a specific topic remains a significant challenge for language models. While standard approaches, such as fine-tuning, are resource-intensive, simpler methods like prompt en…

Prompt Engineering

Marco-LLM: Bridging Languages via Massive Multilingual Training for Cross-Lingual Enhancement

2024-12-05 · Lingfeng Ming, Bo Zeng, Chenyang Lyu, Tianqi Shi 외

Large Language Models (LLMs) have achieved remarkable progress in recent years; however, their excellent performance is still largely limited to major world languages, primarily English. Many LLMs continue to face challe…

BelebeleMachine Translation

Emu: Enhancing Multilingual Sentence Embeddings with Semantic Specialization

2019-09-15 · Wataru Hirota, Yoshihiko Suhara, Behzad Golshan, Wang-Chiew Tan

We present Emu, a system that semantically enhances multilingual sentence embeddings. Our framework fine-tunes pre-trained multilingual sentence embeddings using two main components: a semantic classifier and a language …

intent-classificationIntent ClassificationSemantic SimilaritySemantic Textual Similarity+4

Lens: Rethinking Multilingual Enhancement for Large Language Models

2024-10-06 · Weixiang Zhao, Yulin Hu, Jiahe Guo, Xingyu Sui 외

Despite the growing global demand for large language models (LLMs) that serve users from diverse linguistic backgrounds, most cutting-edge LLMs remain predominantly English-centric. This creates a performance gap across …