paper-with-me

홈 › Papers

OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text

2023-10-10 · Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, Jimmy Ba

There is growing evidence that pretraining on high quality, carefully thought-out tokens such as code or mathematics plays an important role in improving the reasoning abilities of large language models. For example, Minerva, a PaLM model finetuned on billions of tokens of mathematical documents from arXiv and the web, reported dramatically improved performance on problems that require quantitative reasoning. However, because all known open source web datasets employ preprocessing that does not faithfully preserve mathematical notation, the benefits of large scale training on quantitive web documents are unavailable to the research community. We introduce OpenWebMath, an open dataset inspired by these works containing 14.7B tokens of mathematical webpages from Common Crawl. We describe in detail our method for extracting text and LaTeX content and removing boilerplate from HTML documents, as well as our methods for quality filtering and deduplication. Additionally, we run small-scale experiments by training 1.4B parameter language models on OpenWebMath, showing that models trained on 14.7B tokens of our dataset surpass the performance of models trained on over 20x the amount of general language data. We hope that our dataset, openly released on the Hugging Face Hub, will help spur advances in the reasoning abilities of large language models.

📄 PDF Abstract BibTeX arXiv:2310.06786

Code (2)

keirp/OpenWebMath 공식 구현
deepseek-ai/deepseek-math pytorch

Methods 이 논문이 사용한 방법론

PaLM 설명 없음

Similar Papers 제목 키워드 기반

Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset

2025-08-20 · Rabeeh Karimi Mahabadi, Sanjeev Satheesh, Shrimai Prabhumoye, Mostofa Patwary 외 arxiv

Pretraining large language models (LLMs) on high-quality, structured data such as mathematics and code substantially enhances reasoning capabilities. However, existing math-focused datasets built from Common Crawl suffer…

MIND: Math Informed syNthetic Dialogues for Pretraining LLMs

2024-10-15 · Syeda Nahida Akter, Shrimai Prabhumoye, John Kamalu, Sanjeev Satheesh 외

The utility of synthetic data to enhance pretraining data quality and hence to improve downstream task accuracy has been widely explored in recent large language models (LLMs). Yet, these approaches fall inadequate in co…

GSM8KMathMathematical ReasoningMMLU

An Overview of zbMATH Open Digital Library

2024-10-09 · Madhurima Deb, Isabel Beckenbach, Matteo Petrera, Dariush Ehsani 외

Mathematical research thrives on the effective dissemination and discovery of knowledge. zbMATH Open has emerged as a pivotal platform in this landscape, offering a comprehensive repository of mathematical literature. Be…

Information RetrievalRetrieval

AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

2025-04-23 · Ivan Moshkov, Darragh Hanley, Ivan Sorokin, Shubham Toshniwal 외

This paper presents our winning submission to the AI Mathematical Olympiad - Progress Prize 2 (AIMO-2) competition. Our recipe for building state-of-the-art mathematical reasoning models relies on three key pillars. Firs…

MathMathematical Reasoning

Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale

2024-09-25 · Fan Zhou, Zengzhi Wang, Qian Liu, Junlong Li 외

Large language model pre-training has traditionally relied on human experts to craft heuristics for improving the corpora quality, resulting in numerous rules developed to date. However, these rules lack the flexibility …

Large Language Model