paper-with-me

홈 › Papers

Power Lines: Scaling Laws for Weight Decay and Batch Size in LLM Pre-training

2025-05-19 · Shane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray, Daria Soboleva, Joel Hestness

Efficient LLM pre-training requires well-tuned hyperparameters (HPs), including learning rate {\eta} and weight decay {\lambda}. We study scaling laws for HPs: formulas for how to scale HPs as we scale model size N, dataset size D, and batch size B. Recent work suggests the AdamW timescale, B/({\eta}{\lambda}D), should remain constant across training settings, and we verify the implication that optimal {\lambda} scales linearly with B, for a fixed N,D. However, as N,D scale, we show the optimal timescale obeys a precise power law in the tokens-per-parameter ratio, D/N. This law thus provides a method to accurately predict {\lambda}opt in advance of large-scale training. We also study scaling laws for optimal batch size Bopt (the B enabling lowest loss at a given N,D) and critical batch size Bcrit (the B beyond which further data parallelism becomes ineffective). In contrast with prior work, we find both Bopt and Bcrit scale as power laws in D, independent of model size, N. Finally, we analyze how these findings inform the real-world selection of Pareto-optimal N and D under dual training time and compute objectives.

📄 PDF Abstract BibTeX arXiv:2505.13738

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

AdamW AdamW is a stochastic optimization method that modifies the typical implementation of weight decay in Adam, by decoupling [weight…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

Scaling Laws from Sequential Feature Recovery: A Solvable Hierarchical Model

2026-05-14 · Arie Wortsman-Zurich, Hugo Tabanelli, Yatin Dandi, Florent Krzakala 외 arxiv

We propose a simple mechanism by which scaling laws emerge from feature learning in multi-layer networks. We study a high-dimensional hierarchical target that is a globally high-degree function, but that can be represent…

Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

2025-09-29 · Leonardo Defilippis, Yizhou Xu, Julius Girardin, Emanuele Troiani 외 arxiv

Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for…

Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?

2026-03-02 · Jihwan Kim, Dogyoon Song, Chulhee Yun arxiv

We study scaling laws of signSGD under a power-law random features (PLRF) model that accounts for both feature and target decay. We analyze the population risk of a linear model trained with one-pass signSGD on Gaussian-…

Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate Schedules

2025-09-23 · Binghui Li, Fengling Chen, Zixun Huang, Lean Wang 외 arxiv

Scaling laws have emerged as a unifying lens for understanding and guiding the training of large language models (LLMs). However, existing studies predominantly focus on the final-step loss, leaving open whether the enti…

Iterative Orthogonalization Scaling Laws

2025-05-06 · Devan Selvaraj

The muon optimizer has picked up much attention as of late as a possible replacement to the seemingly omnipresent Adam optimizer. Recently, care has been taken to document the scaling laws of hyper-parameters under muon …