paper-with-me

홈 › Papers

HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models

2025-11-02 · Stephan Oepen, Nikolay Arefev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Charpentier, Pinzhen Chen, Mariya Fedorova, Ona de Gibert, Barry Haddow, Jan Hajič, Jindřich Helcl, Andrey Kutuzov, Veronika Laippala, Zihao Li, Risto Luukkonen, Bhavitvya Malik, Vladislav Mikhailov, Amanda Myntti, Dayyán O'Brien, Lucie Poláková, Sampo Pyysalo, Gema Ramírez Sánchez, Janine Siewert, Pavel Stepachev, Jörg Tiedemann, Teemu Vahtola, Dušan Variš, Fedor Vitiugin, Tea Vojtěchová, Jaume Zaragoza arxiv

We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for 24 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder-decoder models, as well as a handful of monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation.

📄 PDF Abstract BibTeX arXiv:2511.01066

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationMachine Translation

Similar Papers 제목 키워드 기반

DHPLT: large-scale multilingual diachronic corpora and word representations for semantic change modelling

2026-02-12 · Mariia Fedorova, Andrey Kutuzov, Khonzoda Umarova arxiv

In this resource paper, we present DHPLT, an open collection of diachronic corpora in 41 diverse languages. DHPLT is based on the web-crawled HPLT datasets; we use web crawl timestamps as the approximate signal of docume…

A New Massive Multilingual Dataset for High-Performance Language Technologies

2024-03-20 · Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón 외

We present the HPLT (High Performance Language Technologies) language resources, a new massive multilingual dataset including both monolingual and bilingual corpora extracted from CommonCrawl and previously unused web cr…

Language ModelingLanguage ModellingMachine TranslationManagement+1

DocHPLT: A Massively Multilingual Document-Level Translation Dataset

2025-08-18 · Dayyán O'Brien, Bhavitvya Malik, Ona de Gibert, Pinzhen Chen 외 arxiv

Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, …

Machine Translation

An Expanded Massive Multilingual Dataset for High-Performance Language Technologies

2025-03-13 · Laurie Burchell, Ona de Gibert, Nikolay Arefyev, Mikko Aulamo 외

Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In this work, we present HPLT v2, a collectio…

Machine TranslationSentence

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

2026-07-01 · Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo, David Salinas 외 arxiv

Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.8 trillion target-language toke…