paper-with-me

홈 › Papers

Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora

2025-09-10 · Thales Sales Almeida, Rodrigo Nogueira, Helio Pedrini arxiv

The performance of large language models (LLMs) is deeply influenced by the quality and composition of their training data. While much of the existing work has centered on English, there remains a gap in understanding how to construct effective training corpora for other languages. We explore scalable methods for building web-based corpora for LLMs. We apply them to build a new 120B token corpus in Portuguese that achieves competitive results to an industrial-grade corpus. Using a continual pretraining setup, we study how different data selection and preprocessing strategies affect LLM performance when transitioning a model originally trained in English to another language. Our findings demonstrate the value of language-specific filtering pipelines, including classifiers for education, science, technology, engineering, and mathematics (STEM), as well as toxic content. We show that adapting a model to the target language leads to performance improvements, reinforcing the importance of high-quality, language-specific data. While our case study focuses on Portuguese, our methods are applicable to other languages, offering insights for multilingual LLM development.

📄 PDF Abstract BibTeX arXiv:2509.08824

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Pretraining

Similar Papers 제목 키워드 기반

NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus

2026-04-30 · Enzo S. N. Silva, Pablo B. Costa, Raphael C. Vlasman, Rosimeire P. Costa 외 arxiv

High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTimbau and Albertina PT-BR, we introduce NorBERTo, a modern encoder ba…

Semantic Similarity

Tucano 2 Cool: Better Open Source LLMs for Portuguese

2026-03-03 · Nicholas Kluge Corrêa, Aniket Sen, Shiza Fatimah, Sophia Falk 외 arxiv

We present Tucano 2, a fully open suite of large language models (LLMs) with 0.5-3.7 billion parameters, designed to address certain gaps in open-source development for Portuguese LLMs. Following our previous works, we n…

Continual Pretraining

Introducing Bode: A Fine-Tuned Large Language Model for Portuguese Prompt-Based Task

2024-01-05 · Gabriel Lino Garcia, Pedro Henrique Paiola, Luis Henrique Morelli, Giovani Candido 외

Large Language Models (LLMs) are increasingly bringing advances to Natural Language Processing. However, low-resource languages, those lacking extensive prominence in datasets for various NLP tasks, or where existing dat…

In-Context LearningLanguage ModelingLanguage ModellingLarge Language Model

AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese

2026-03-27 · Afonso Simplício, Gonçalo Vinagre, Miguel Moura Ramos, Diogo Tavares 외 arxiv

Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translated benchmarks likely missing the variant…

Building The First English-Brazilian Portuguese Corpus for Automatic Post-Editing

2020-12-01 · COLING 2020 8 · Felipe Almeida Costa, Thiago castro Ferreira, Adriana Pagano, Wagner Meira

This paper introduces the first corpus for Automatic Post-Editing of English and a low-resource language, Brazilian Portuguese. The source English texts were extracted from the WebNLG corpus and automatically translated …

Automatic Post-EditingMachine TranslationTranslation