paper-with-me

홈 › Papers

FinerWeb-10BT: Refining Web Data with LLM-Based Line-Level Filtering

2025-01-13 · Erik Henriksson, Otto Tarkka, Filip Ginter

Data quality is crucial for training Large Language Models (LLMs). Traditional heuristic filters often miss low-quality text or mistakenly remove valuable content. In this paper, we introduce an LLM-based line-level filtering method to enhance training data quality. We use GPT-4o mini to label a 20,000-document sample from FineWeb at the line level, allowing the model to create descriptive labels for low-quality lines. These labels are grouped into nine main categories, and we train a DeBERTa-v3 classifier to scale the filtering to a 10B-token subset of FineWeb. To test the impact of our filtering, we train GPT-2 models on both the original and the filtered datasets. The results show that models trained on the filtered data achieve higher accuracy on the HellaSwag benchmark and reach their performance targets faster, even with up to 25\% less data. This demonstrates that LLM-based line-level filtering can significantly improve data quality and training efficiency for LLMs. We release our quality-annotated dataset, FinerWeb-10BT, and the codebase to support further work in this area.

📄 PDF Abstract BibTeX arXiv:2501.07314

Code (1)

turkunlp/finerweb-10bt 공식 구현 pytorch

Tasks

DescriptiveHellaSwag

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Adam 설명 없음
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Linear Warmup With Cosine Annealing Linear Warmup With Cosine Annealing is a learning rate schedule where we increase the learning rate linearly for $n$ updates and then anneal according to a cosine schedule…

Similar Papers 제목 키워드 기반

FiNERweb: Datasets and Artifacts for Scalable Multilingual Named Entity Recognition

2025-12-15 · Jonas Golde, Patrick Haller, Alan Akbik arxiv

Recent multilingual named entity recognition (NER) work has shown that large language models (LLMs) can provide effective synthetic supervision, yet such datasets have mostly appeared as by-products of broader experiment…

Unified Unsupervised Anomaly Detection via Matching Cost Filtering

2025-10-03 · Zhe Zhang, Mingxiu Cai, Gaochang Wu, Jing Zhang 외 arxiv

Unsupervised anomaly detection (UAD) aims to identify image- and pixel-level anomalies using only normal training data, with wide applications such as industrial inspection and medical analysis, where anomalies are scarc…

Unsupervised Anomaly Detection

Coarse2Fine: Two-Layer Fusion For Image Retrieval

2016-07-04 · Gaipeng Kong, Le Dong, Wenpu Dong, Liang Zheng 외

This paper addresses the problem of large-scale image retrieval. We propose a two-layer fusion method which takes advantage of global and local cues and ranks database images from coarse to fine (C2F). Departing from the…

Image RetrievalRetrievalVocal Bursts Valence Prediction

Bilevel Data Curation for LLM Fine-tuning: Offline Selection and Online Self-Refining Generation

2025-11-26 · Quan Xiao, Yutong Xuan, Gaowen Liu, Ramana Rao Kompella 외 arxiv

Supervised fine-tuning (SFT) datasets are critical to the downstream performance of large language models, yet they often contain low-quality or harmful question-response pairs. To improve SFT data quality, we develop a …

Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection

2026-04-22 · Yassine Turki, Vinko Sabolčec, Bettina Messmer, Martin Jaggi arxiv

As Large Language Models (LLMs) scale, data curation has shifted from maximizing volume to optimizing the signal-to-noise ratio by performing quality filtering. However, for many languages, native high quality data is in…

Cross-Lingual Transfer