paper-with-me

홈 › Papers

MegaMath: Pushing the Limits of Open Math Corpora

2025-04-03 · Fan Zhou, Zengzhi Wang, Nikhil Ranjan, Zhoujun Cheng, Liping Tang, Guowei He, Zhengzhong Liu, Eric P. Xing

Mathematical reasoning is a cornerstone of human intelligence and a key benchmark for advanced capabilities in large language models (LLMs). However, the research community still lacks an open, large-scale, high-quality corpus tailored to the demands of math-centric LLM pre-training. We present MegaMath, an open dataset curated from diverse, math-focused sources through following practices: (1) Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on the Internet. (2) Recalling Math-related code data: We identified high quality math-related code from large code training corpus, Stack-V2, further enhancing data diversity. (3) Exploring Synthetic data: We synthesized QA-style text, math-related code, and interleaved text-code blocks from web data or code data. By integrating these strategies and validating their effectiveness through extensive ablations, MegaMath delivers 371B tokens with the largest quantity and top quality among existing open math pre-training datasets.

📄 PDF Abstract BibTeX arXiv:2504.02807

Code (1)

llm360/megamath 공식 구현

Tasks

DiversityMathMathematical Reasoning

Similar Papers 제목 키워드 기반

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

2025-06-25 · Zengzhi Wang, Fan Zhou, Xuefeng Li, PengFei Liu

Different base language model families, such as Llama and Qwen, exhibit divergent behaviors during post-training with reinforcement learning (RL), especially on reasoning-intensive tasks. What makes a base language model…

Language ModelingLanguage ModellingMathreinforcement-learning+2

Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset

2025-08-20 · Rabeeh Karimi Mahabadi, Sanjeev Satheesh, Shrimai Prabhumoye, Mostofa Patwary 외 arxiv

Pretraining large language models (LLMs) on high-quality, structured data such as mathematics and code substantially enhances reasoning capabilities. However, existing math-focused datasets built from Common Crawl suffer…

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

2024-02-05 · Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu 외

Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B wit…

Arithmetic ReasoningMathMathematical ReasoningMath Word Problem Solving

Physics-Driven Deep Learning for Computational Magnetic Resonance Imaging

2022-03-23 · Kerstin Hammernik, Thomas Küstner, Burhaneddin Yaman, Zhengnan Huang 외

Physics-driven deep learning methods have emerged as a powerful tool for computational magnetic resonance imaging (MRI) problems, pushing reconstruction performance to new limits. This article provides an overview of the…

Deep LearningMRI Reconstruction

Every Sample a Task: Pushing the Limits of Heterogeneous Models with Personalized Regression

2019-05-16 · ICML Workshop AMTL 2019 6 · Anonymous

When data arise from multiple latent subpopulations, machine learning frameworks typically estimate parameter values independently for each sub-population. In this paper, we propose to overcome these limits by considerin…

BIG-bench Machine Learningregression