paper-with-me

홈 › Papers

Scaling Domain Data Repetition in LLM Pretraining

2026-08-14 · Jingwei Li, Xinran Gu, Rui Dai, Xintong Hao, Chengyin Xu, Yan Wu, Shuran Zheng, Jingzhao Zhang arxiv

As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(TPP\)). However, high-quality domain data is much harder to scale than general web data. As model size and the training-token budget increase, its fraction in the training mixture tends to decrease. Repeating the available high-quality data provides an effective way to counteract this dilution, but excessive repetition may lead to overfitting. We study this trade-off under practical LLM scaling, where the training-token budget grows proportionally with model size. For a fixed domain, we first find that, surprisingly at a fixed \(TPP\), the optimal repetition count mildly increases with model size. Across different domains, we find that the optimal repetition count is strongly negatively correlated with the final validation loss of a domain: domains with lower loss can generally benefit from more repetitions. In contrast, the amount of unique domain data is only weakly related to the optimal repetition count. These findings suggest that repetition counts tuned on smaller proxy models with the same \(TPP\) can provide a practical estimate for larger models.

📄 PDF Abstract BibTeX arXiv:2608.14071

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scaling Laws for Mixture Pretraining Under Data Constraints

2026-05-12 · Anastasiia Sedova, Skyler Seto, Natalie Schluter, Pierre Ablin arxiv

As language models scale, the amount of data they require grows -- yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this sca…

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition

2026-05-04 · Fengze Liu, Weidong Zhou, Binbin Liu, Ping Guo 외 arxiv

Upweighting high-quality data in LLM pretraining often improves performance, but in datalimited regimes, especially under overtraining, stronger upweighting increases repetition and can degrade performance. However, stan…

Bridging Compute- and Data-Optimal Pretraining

2026-07-28 · Tian Qin, Kimia Hamidieh, David Alvarez-Melis arxiv

Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We …

Generating Pretraining Tokens from Organic Data for Data-Bound Scaling

2026-05-18 · Zichun Yu, Chenyan Xiong arxiv

LLM pretraining is shifting from a compute-bound to a data-bound regime, where available human (organic) text falls far short of scaling demands. However, reaching the data-bound regime does not mean the model has fully …

Synthetic Data GenerationReinforcement Learning

Reformulation for Pretraining Data Augmentation

2025-02-06 · Xintong Hao, Ruijie Zhu, Ge Zhang, Ke Shen 외

Despite the impressive capabilities of large language models across various tasks, their continued scaling is severely hampered not only by data scarcity but also by the performance degradation associated with excessive …

Data AugmentationPrompt Engineering