paper-with-me

홈 › Papers

Multilingual Language Model Pretraining using Machine-translated Data

2025-02-18 · Jiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin, David Adelani, Yihong Chen, Raphael Tang, Pontus Stenetorp

High-resource languages such as English, enables the pretraining of high-quality large language models (LLMs). The same can not be said for most other languages as LLMs still underperform for non-English languages, likely due to a gap in the quality and diversity of the available multilingual pretraining corpora. In this work, we find that machine-translated texts from a single high-quality source language can contribute significantly to the pretraining quality of multilingual LLMs. We translate FineWeb-Edu, a high-quality English web dataset, into nine languages, resulting in a 1.7-trillion-token dataset, which we call TransWebEdu and pretrain a 1.3B-parameter model, TransWebLLM, from scratch on this dataset. Across nine non-English reasoning tasks, we show that TransWebLLM matches or outperforms state-of-the-art multilingual models trained using closed data, such as Llama3.2, Qwen2.5, and Gemma, despite using an order of magnitude less data. We demonstrate that adding less than 5% of TransWebEdu as domain-specific pretraining data sets a new state-of-the-art in Arabic, Italian, Indonesian, Swahili, and Welsh understanding and commonsense reasoning tasks. To promote reproducibility, we release our corpus, models, and training pipeline under Open Source Initiative-approved licenses.

📄 PDF Abstract BibTeX arXiv:2502.13252

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Multilingual Multimodal Learning with Machine Translated Text

2022-10-24 · Chen Qiu, Dan Oneata, Emanuele Bugliarello, Stella Frank 외

Most vision-and-language pretraining research focuses on English tasks. However, the creation of multilingual multimodal evaluation datasets (e.g. Multi30K, xGQA, XVNLI, and MaRVL) poses a new challenge in finding high-q…

Zero-Shot Cross-Lingual Image-to-Text RetrievalZero-Shot Cross-Lingual Text-to-Image RetrievalZero-Shot Cross-Lingual Visual Natural Language InferenceZero-Shot Cross-Lingual Visual Question Answering+1

Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language

2024-10-31 · Jiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin 외

English, as a very high-resource language, enables the pretraining of high-quality large language models (LLMs). The same cannot be said for most other languages, as leading LLMs still underperform for non-English langua…

Crosslingual Generalization through Multitask Finetuning

2022-11-03 · Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts 외

Multitask prompted finetuning (MTF) has been shown to help large language models generalize to new tasks in a zero-shot setting, but so far explorations of MTF have focused on English data and models. We apply MTF to the…

Coreference ResolutionCross-Lingual TransferQuestion AnsweringSentence Completion+2

DOCmT5: Document-Level Pretraining of Multilingual Language Models

2021-12-16 · Findings (NAACL) 2022 7 · Chia-Hsuan Lee, Aditya Siddhant, Viresh Ratnakar, Melvin Johnson

In this paper, we introduce DOCmT5, a multilingual sequence-to-sequence language model pretrained with large scale parallel documents. While previous approaches have focused on leveraging sentence-level parallel data, we…

de-enDocument SummarizationDocument TranslationLanguage Modeling+4

Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text

2025-09-22 · Dan John Velasco, Matthew Theodore Roque arxiv

Most languages lack sufficient data for large-scale monolingual pretraining, creating a "data wall." Multilingual pretraining helps but is limited by language imbalance and the "curse of multilinguality." An alternative …

Machine Translation