paper-with-me

Papers

Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking

2025-09-25 · Yuandong Tian arxiv

While the phenomenon of grokking, i.e., delayed generalization, has been studied extensively, it remains an open problem whether there is a mathematical framework that characterizes what kind of features will emerge, how and in which conditions it happens, and is closely related to the gradient dynamics of the training, for complex structured inputs. We propose a novel framework, named $\mathbf{Li}_2$, that captures three key stages for the grokking behavior of 2-layer nonlinear networks: (I) Lazy learning, (II) independent feature learning and (III) interactive feature learning. At the lazy learning stage, top layer overfits to random hidden representation and the model appears to memorize, and at the same time, the backpropagated gradient $G_F$ from the top layer now carries information about the target label, with a specific structure that enables each hidden node to learn their representation independently. Interestingly, the independent dynamics follows exactly the gradient ascent of an energy function $E$, and its local maxima are precisely the emerging features. We study whether these local-optima induced features are generalizable, their representation power, and how they change on sample size, in group arithmetic tasks. When hidden nodes start to interact in the later stage of learning, we provably show how $G_F$ changes to focus on missing features that need to be learned. Our study sheds lights on roles played by key hyperparameters such as weight decay, learning rate and sample sizes in grokking, leads to provable scaling laws of feature emergence, memorization and generalization, and reveals why recent optimizers such as Muon can be effective, from the first principles of gradient dynamics. Our analysis can be extended to multi-layers. The code is available at https://github.com/yuandong-tian/understanding/tree/main/ssl/real-dataset/cogo.

📄 PDF Abstract BibTeX arXiv:2509.21519

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Universal scaling laws rule explosive growth inhuman cancers

2021-04-22 · Víctor M. Pérez-García, Gabriel F. Calvo, Jesús J. Bosque, Odelaisy León-Triana 외

Most physical and other natural systems are complex entities composed of a large number of interacting individual elements. It is a surprising fact that they often obey the so-called scaling laws relating an observable q…

Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks

2025-05-19 · Francesco D'Amico, Dario Bocchi, Matteo Negri

Scaling laws in deep learning - empirical power-law relationships linking model performance to resource growth - have emerged as simple yet striking regularities across architectures, datasets, and tasks. These laws are …

Scaling Laws Are Unreliable for Downstream Tasks: A Reality Check

2025-07-01 · Nicholas Lourie, Michael Y. Hu, Kyunghyun Cho

Downstream scaling laws aim to predict task performance at larger scales from pretraining losses at smaller scales. Whether this prediction should be possible is unclear: some works demonstrate that task performance foll…

Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime

2025-09-29 · Leonardo Defilippis, Yizhou Xu, Julius Girardin, Emanuele Troiani 외 arxiv

Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for…

Scaling Laws in Jet Classification

2023-12-04 · Joshua Batson, Yonatan Kahn

We demonstrate the emergence of scaling laws in the benchmark top versus QCD jet classification problem in collider physics. Six distinct physically-motivated classifiers exhibit power-law scaling of the binary cross-ent…

Classification