paper-with-me

Papers

Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering

2025-11-02 · Vlad Negoita, Mihai Masala, Traian Rebedea arxiv

Large Language Models (LLMs) have recently exploded in popularity, often matching or outperforming human abilities on many tasks. One of the key factors in training LLMs is the availability and curation of high-quality data. Data quality is especially crucial for under-represented languages, where high-quality corpora are scarce. In this work we study the characteristics and coverage of Romanian pretraining corpora and we examine how they differ from English data. By training a lightweight multitask model on carefully LLM-annotated Romanian texts, we are able to analyze and perform multi-level filtering (e.g., educational value, topic, format) to generate high-quality pretraining datasets. Our experiments show noteworthy trends in the topics present in Romanian and English data, while also proving the effectiveness of filtering data through improved LLM pretraining performance across multiple benchmarks.

📄 PDF Abstract BibTeX arXiv:2511.01090

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FuLG: 150B Romanian Corpus for Language Model Pretraining

2024-07-18 · Vlad-Andrei Bădoiu, Mihai-Valentin Dumitru, Alexandru M. Gherghescu, Alexandru Agache 외

Research in the field of language models is rapidly evolving, with many open models being released to the public. Openly available pretraining corpora usually focus on only a handful of languages, with many others either…

Language ModelingLanguage Modellingmodel

TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction

2026-01-15 · Mihai Dan Nadas, Laura Diosan, Andreea Tomescu, Andrei Piscoran arxiv

Recent advances in synthetic data generation have shown that compact language models can be trained effectively when the underlying corpus is structurally controlled and linguistically coherent. However, for morphologica…

Synthetic Data GenerationKnowledge Distillation

LLMic: Romanian Foundation Language Model

2025-01-13 · Vlad-Andrei Bădoiu, Mihai-Valentin Dumitru, Alexandru M. Gherghescu, Alexandru Agache 외

Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks with commercial models leading the way. While open models usually operate at a smaller scale, they maintain c…

Language ModelingLanguage ModellingmodelTranslation

Neural Grammatical Error Correction for Romanian

2026-04-26 · Teodor-Mihai Cotet, Stefan Ruseti, Mihai Dascalu arxiv

Resources for Grammatical Error Correction (GEC) in non-English languages are scarce, while available spellcheckers in these languages are mostly limited to simple corrections and rules. In this paper we introduce a firs…

Grammatical Error Correction

"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models

2026-05-29 · Mihai Masala, Marius Leordeanu, Mihai Dascalu, Traian Rebedea arxiv

Vision-Language Models (VLMs) largely follow the text-only LLM trajectory, excelling on English benchmarks but sharply degrading on low-resource languages, where neither large-scale image-text corpora nor culturally grou…

Machine TranslationVisual Grounding