paper-with-me

홈 › Papers

Reusing Overtrained Language Models Saturates Scaling

2025-10-08 · Seng Pei Liew, Takuya Kato arxiv

Reusing pretrained base models for further pretraining, such as continual pretraining or model growth, is promising at reducing the cost of training language models from scratch. However, the effectiveness remains unclear, especially when applied to overtrained base models. In this work, we empirically study the scaling properties of model reuse and find that the scaling efficiency diminishes in a predictable manner: The scaling exponent with respect to second-stage training tokens decreases logarithmically with the number of tokens used to pretrain the base model. The joint dependence on first- and second-stage tokens is accurately modeled by a simple scaling law. Such saturation effect reveals a fundamental trade-off in multi-stage pretraining strategies: the more extensively a base model is pretrained, the less benefit additional pretraining provides. Our findings provide practical insights for efficient language model training and raise important considerations for the reuse of overtrained models.

📄 PDF Abstract BibTeX arXiv:2510.06548

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Pretraining

Similar Papers 제목 키워드 기반

Will we run out of data? Limits of LLM scaling based on human-generated data

2022-10-26 · Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu 외

We investigate the potential constraints on LLM scaling posed by the availability of public human-generated text data. We forecast the growing demand for training data based on current trends and estimate the total stock…

Language ModelingLanguage ModellingSynthetic Data GenerationTransfer Learning

Establishing Task Scaling Laws via Compute-Efficient Model Ladders

2024-12-05 · Akshita Bhagia, Jiacheng Liu, Alexander Wettig, David Heineman 외

We develop task scaling laws and model ladders to predict the individual task performance of pretrained language models (LMs) in the overtrained setting. Standard power laws for language modeling loss cannot accurately m…

Language ModelingLanguage ModellingMultiple-choice

Test-Time Scaling Makes Overtraining Compute-Optimal

2026-04-01 · Nicholas Roberts, Sungjun Cho, Zhiqi Gao, Tzu-Heng Huang 외 arxiv

Modern LLMs scale at test-time, e.g. via repeated sampling, where inference cost grows with model size and the number of samples. This creates a trade-off that pretraining scaling laws, such as Chinchilla, do not address…

Mechanistic Design and Scaling of Hybrid Architectures

2024-03-26 · Michael Poli, Armin W Thomas, Eric Nguyen, Pragaash Ponnusamy 외

The development of deep learning architectures is a resource-demanding process, due to a vast design space, long prototyping times, and high compute costs associated with at-scale model training and evaluation. We set ou…

Mamba

Quantifying Overfitting: Evaluating Neural Network Performance through Analysis of Null Space

2023-05-30 · Hossein Rezaei, Mohammad Sabokrou

Machine learning models that are overfitted/overtrained are more vulnerable to knowledge leakage, which poses a risk to privacy. Suppose we download or receive a model from a third-party collaborator without knowing its …