paper-with-me

홈 › Papers

On the Optimizer Dependence of Neural Scaling Laws

2026-05-28 · Vansh Ramani, Shourya Vir Jain arxiv

The scaling exponent $α$ in neural scaling laws $L(N) \propto N^{-α}$ is commonly treated as a fixed constant set by architecture and data. We present evidence that $α$ depends systematically on the optimizer. In controlled random-feature regression experiments -- the canonical theoretical framework for neural scaling -- we measure $α$ across five optimizer variants and six spectral conditions. Preconditioned optimizers consistently yield steeper scaling (larger $α$), with the $α$-shift increasing across most of the tested spectral range, peaking near $s = 1.5$, and remaining large at $s = 2.0$. At $s \approx 1.0$ (characteristic of natural language), the full natural gradient achieves $α\approx 0.31$ versus $α\approx 0.12$ for gradient descent -- a $2.6\times$ larger fitted exponent that, within the random-feature model, compounds with each model-size doubling. Whether and how this exponent shift transfers to large-scale LLM training -- where recent evidence suggests the advantage may attenuate with scale -- remains an important open question. Our results imply that scaling-law forecasts should account for optimizer choice, and we provide a spectral diagnostic predicting when advanced optimizers will pay off.

📄 PDF Abstract BibTeX arXiv:2605.29387

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Robust Scaling Laws for Optimizers

2026-02-07 · Alexandra Volkova, Mher Safaryan, Christoph H. Lampert, Dan Alistarh arxiv

The quality of Large Language Model (LLM) pretraining depends on multiple factors, including the compute budget and the choice of optimization algorithm. Empirical scaling laws are widely used to predict loss as model si…

Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws

2026-05-20 · Nandan Kumar Jha, Brandon Reagen arxiv

Scaling laws have made language-model performance predictable from model size, data, and compute, but they typically treat the optimizer as a fixed training detail. We show that this assumption misses a fundamental axis …

Iterative Orthogonalization Scaling Laws

2025-05-06 · Devan Selvaraj

The muon optimizer has picked up much attention as of late as a possible replacement to the seemingly omnipresent Adam optimizer. Recently, care has been taken to document the scaling laws of hyper-parameters under muon …

Scaling laws for amplitude surrogates

2026-01-19 · Henning Bahl, Victor Bresó-Pla, Anja Butter, Joaquín Iturriza Ramirez arxiv

Scaling laws describing the dependence of neural network performance on the amount of training data, the spent compute, and the network size have emerged across a huge variety of machine learning task and datasets. In th…

Resolving Discrepancies in Compute-Optimal Scaling of Language Models

2024-06-27 · Tomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt 외

Kaplan et al. and Hoffmann et al. developed influential scaling laws for the optimal model size as a function of the compute budget, but these laws yield substantially different predictions. We explain the discrepancy by…