paper-with-me

Papers

FuLG: 150B Romanian Corpus for Language Model Pretraining

2024-07-18 · Vlad-Andrei Bădoiu, Mihai-Valentin Dumitru, Alexandru M. Gherghescu, Alexandru Agache, Costin Raiciu

Research in the field of language models is rapidly evolving, with many open models being released to the public. Openly available pretraining corpora usually focus on only a handful of languages, with many others either missing completely or extremely underrepresented. In this report, we introduce FuLG, a hundred-fifty-billion-token Romanian corpus extracted from CommonCrawl. We present our methodology for filtering FuLG and compare it via ablation studies against existing Romanian corpora.

📄 PDF Abstract BibTeX arXiv:2407.13657

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modellingmodel

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Neural Grammatical Error Correction for Romanian

2026-04-26 · Teodor-Mihai Cotet, Stefan Ruseti, Mihai Dascalu arxiv

Resources for Grammatical Error Correction (GEC) in non-English languages are scarce, while available spellcheckers in these languages are mostly limited to simple corrections and rules. In this paper we introduce a firs…

Grammatical Error Correction

LLMic: Romanian Foundation Language Model

2025-01-13 · Vlad-Andrei Bădoiu, Mihai-Valentin Dumitru, Alexandru M. Gherghescu, Alexandru Agache 외

Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks with commercial models leading the way. While open models usually operate at a smaller scale, they maintain c…

Language ModelingLanguage ModellingmodelTranslation

TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction

2026-01-15 · Mihai Dan Nadas, Laura Diosan, Andreea Tomescu, Andrei Piscoran arxiv

Recent advances in synthetic data generation have shown that compact language models can be trained effectively when the underlying corpus is structurally controlled and linguistically coherent. However, for morphologica…

Synthetic Data GenerationKnowledge Distillation

Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering

2025-11-02 · Vlad Negoita, Mihai Masala, Traian Rebedea arxiv

Large Language Models (LLMs) have recently exploded in popularity, often matching or outperforming human abilities on many tasks. One of the key factors in training LLMs is the availability and curation of high-quality d…

Adapting the TTL Romanian POS Tagger to the Biomedical Domain

2017-09-01 · RANLP 2017 9 · Maria Mitrofan, Radu Ion

This paper presents the adaptation of the Hidden Markov Models-based TTL part-of-speech tagger to the biomedical domain. TTL is a text processing platform that performs sentence splitting, tokenization, POS tagging, chun…

ChunkingDomain AdaptationLemmatizationnamed-entity-recognition+8