paper-with-me

홈 › Papers

Generating Pretraining Tokens from Organic Data for Data-Bound Scaling

2026-05-18 · Zichun Yu, Chenyan Xiong arxiv

LLM pretraining is shifting from a compute-bound to a data-bound regime, where available human (organic) text falls far short of scaling demands. However, reaching the data-bound regime does not mean the model has fully utilized its organic corpus. In this paper, we introduce SynPro, a synthetic data generation framework that helps LLMs more thoroughly learn from limited organic data. SynPro applies two operations, rephrasing and reformat, that present the same organic source in diverse forms to facilitate deeper learning without introducing external information. Both generators are optimized via reinforcement learning with quality, faithfulness, and data influence rewards, and are continuously updated as pretraining plateaus to target content the model has yet to absorb. We pretrain 400M and 1.1B models with 10% of their Chinchilla-optimal tokens (0.8B and 2.2B) from DCLM-Baseline, reflecting a realistic data-bound regime in frontier pretraining. Our results reveal that organic data is significantly underutilized by standard repetition: SynPro unlocks 3.7-5.2x the effective tokens of repetition, even surpassing the non-data-bound oracle that trains on equivalent unique data at the 1.1B scale. Analyses confirm that faithful, model-aware synthesis sustains data-bound scaling without causing distribution collapse. We open-source our code at https://github.com/cxcscmu/SynPro.

📄 PDF Abstract BibTeX arXiv:2605.17849

Code (0)

등록된 구현이 없습니다.

Tasks

Synthetic Data GenerationReinforcement Learning

Similar Papers 제목 키워드 기반

RePro: Training Language Models to Faithfully Recycle the Web for Pretraining

2025-10-12 · Zichun Yu, Chenyan Xiong arxiv

High-quality pretraining data is the fossil fuel of large language models (LLMs), yet its reserves are running low for frontier models. In this paper, we introduce RePro, a novel web recycling method that trains a relati…

Reinforcement Learning

Learning to Commit: Generating Organic Pull Requests via Online Repository Memory

2026-03-27 · Mo Li, L. H. Xu, Qitai Tan, Ting Cao 외 arxiv

Large language model (LLM)-based coding agents achieve impressive results on controlled benchmarks yet routinely produce pull requests that real maintainers reject. The root cause is not functional incorrectness but a la…

Linking the Neural Machine Translation and the Prediction of Organic Chemistry Reactions

2016-12-29 · Juno Nam, Jurae Kim

Finding the main product of a chemical reaction is one of the important problems of organic chemistry. This paper describes a method of applying a neural machine translation model to the prediction of organic chemical re…

Machine TranslationTranslation

Specialising and Analysing Instruction-Tuned and Byte-Level Language Models for Organic Reaction Prediction

2024-05-17 · Jiayun Pang, Ivan Vulić

Transformer-based encoder-decoder models have demonstrated impressive results in chemical reaction prediction tasks. However, these models typically rely on pretraining using tens of millions of unlabelled molecules, whi…

Chemical Reaction PredictionDecoderGPUPrediction

Bridging the Gap between Chemical Reaction Pretraining and Conditional Molecule Generation with a Unified Model

2023-03-13 · Bo Qiang, Yiran Zhou, Yuheng Ding, Ningfeng Liu 외

Chemical reactions are the fundamental building blocks of drug design and organic chemistry research. In recent years, there has been a growing need for a large-scale deep-learning framework that can efficiently capture …

Drug DesignRepresentation Learning