paper-with-me

Papers

Blu-WERP (Web Extraction and Refinement Pipeline): A Scalable Pipeline for Preprocessing Large Language Model Datasets

2025-11-22 · Gowtham, Sai Rupesh, Sanjay Kumar, Saravanan, Venkata Chaithanya arxiv

High-quality training data is fundamental to large language model (LLM) performance, yet existing preprocessing pipelines often struggle to effectively remove noise and unstructured content from web-scale corpora. This paper presents Blu-WERP, a novel data preprocessing pipeline designed to optimize the quality of Common Crawl WARC files for LLM training. We demonstrate that Blu-WERP significantly outperforms established baselines including DCLM across multiple model scales and evaluation benchmarks. Our pipeline processes CC WARC dumps, implementing advanced filtering and quality assessment mechanisms. We conducted comprehensive evaluations using models with 150M, 400M, 530M, 750M, and 1B parameters, testing against nine standard benchmarks categorized as World Knowledge & Reasoning, Language Understanding, and Commonsense Reasoning. Results show Blu-WERP consistently achieved superior performance across all model scales. At the 1B parameter scale, Relatively Blu-WERP demonstrates a 4.0% and 9.5% aggregate improvement over DCLM and Fineweb respectively, while achieving quality-per-token efficiency gain. Categorical analysis reveals 2.4% improvement in World Knowledge & Reasoning, 6.2% improvement in Language Understanding, and 4.2% improvement in Commonsense Reasoning. These results establish Blu-WERP as a state-of-the-art preprocessing pipeline that substantially improves LLM training data quality and downstream model performance with reduced computational cost. Our findings contribute to the growing body of research on data-centric AI, demonstrating that preprocessing pipeline design significantly impacts LLM capabilities. The Blu-WERP pipeline represents a practical advancement in data quality optimization, offering researchers and practitioners an effective solution for improving LLM training efficiency and model performance.

📄 PDF Abstract BibTeX arXiv:2511.18054

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT

2025-10-24 · Hyeonsu Kang, Emily Bao, Anjan Goswami arxiv

Vision-language models (VLMs) are increasingly used to evaluate multimodal content, including presentation slides, yet their slide-specific understanding remains underexplored {despite their growing role as critics in ag…

Scalable Robust Graph and Feature Extraction for Arbitrary Vessel Networks in Large Volumetric Datasets

2021-02-05 · Dominik Drees, Aaron Scherzinger, René Hägerling, Friedemann Kiefer 외

Recent advances in 3D imaging technologies provide novel insights to researchers and reveal finer and more detail of examined specimen, especially in the biomedical domain, but also impose huge challenges regarding scala…

Foreground Segmentation

Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text

2026-01-15 · Zhihao Xu, Rumei Li, Jiahuan Li, Rongxiang Weng 외 arxiv

Enabling Large Language Models (LLMs) to effectively utilize tools in multi-turn interactions is essential for building capable autonomous agents. However, acquiring diverse and realistic multi-turn tool-use data remains…

UltraShape 1.0: High-Fidelity 3D Shape Generation via Scalable Geometric Refinement

2025-12-24 · Tanghui Jia, Dongyu Yan, Dehao Hao, Yang Li 외 arxiv

In this report, we introduce UltraShape 1.0, a scalable 3D diffusion framework for high-fidelity 3D geometry generation. The proposed approach adopts a two-stage generation pipeline: a coarse global structure is first sy…

3D Generation

DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering

2025-11-30 · Toshiki Katsube, Taiga Fukuhara, Kenichiro Ando, Yusuke Mukuta 외 arxiv

This work addresses the scarcity of high-quality, large-scale resources for Japanese Vision-and-Language (V&L) modeling. We present a scalable and reproducible pipeline that integrates large-scale web collection with rig…

Visual Question AnsweringImage Captioning