paper-with-me

홈 › Papers

An Expanded Massive Multilingual Dataset for High-Performance Language Technologies

2025-03-13 · Laurie Burchell, Ona de Gibert, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, and Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajič, and Jindřich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, and Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Petter Mæhlum, and Bhavitvya Malik, Farrokh Mehryary, Vladislav Mikhailov, Nikita Moghe, and Amanda Myntti, Dayyán O'Brien, Stephan Oepen, Proyag Pal, Jousia Piha, and Sampo Pyysalo, Gema Ramírez-Sánchez, David Samuel, Pavel Stepachev, and Jörg Tiedemann, Dušan Variš, Tereza Vojtěchová, Jaume Zaragoza-Bernabeu

Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In this work, we present HPLT v2, a collection of high-quality multilingual monolingual and parallel corpora. The monolingual portion of the data contains 8T tokens covering 193 languages, while the parallel data contains 380M sentence pairs covering 51 languages. We document the entire data pipeline and release the code to reproduce it. We provide extensive analysis of the quality and characteristics of our data. Finally, we evaluate the performance of language models and machine translation systems trained on HPLT v2, demonstrating its value.

📄 PDF Abstract BibTeX arXiv:2503.10267

Code (1)

hplt-project/hplt-textpipes 공식 구현

Tasks

Machine TranslationSentence

Similar Papers 제목 키워드 기반

Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark

2026-04-14 · Terra Blevins, Stephen Mayhew, Marek Šuppa, Hila Gonen 외 arxiv

While multilingual language models promise to bring the benefits of LLMs to speakers of many languages, gold-standard evaluation benchmarks in most languages to interrogate these assumptions remain scarce. The Universal …

Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation

2025-04-09 · Israfel Salazar, Manuel Fernández Burda, Shayekh Bin Islam, Arshia Soltani Moakhar 외

The evaluation of vision-language models (VLMs) has mainly relied on English-language benchmarks, leaving significant gaps in both multilingual and multicultural coverage. While multilingual benchmarks have expanded, bot…

Multiple-choice

Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond

2024-08-07 · Beomseok Lee, Ioan Calapodescu, Marco Gaido, Matteo Negri 외

We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different famil…

BenchmarkingLanguage Identificationslot-fillingSlot Filling+1

Assessment of Massively Multilingual Sentiment Classifiers

2022-04-11 · WASSA (ACL) 2022 5 · Krzysztof Rajda, Łukasz Augustyniak, Piotr Gramacki, Marcin Gruza 외

Models are increasing in size and complexity in the hunt for SOTA. But what if those 2\% increase in performance does not make a difference in a production use case? Maybe benefits from a smaller, faster model outweigh t…

Sentiment Analysis

One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers

2025-06-12 · Diana Abagyan, Alejandro R. Salamanca, Andres Felipe Cruz-Salinas, Kris Cao 외

Pretraining massively multilingual Large Language Models (LLMs) for many languages at once is challenging due to limited model capacity, scarce high-quality data, and compute constraints. Moreover, the lack of language c…

All