paper-with-me

홈 › Papers

Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures

2025-12-23 · Yedi Zhang, Andrew Saxe, Peter E. Latham arxiv

Neural networks trained with gradient descent often learn solutions of increasing complexity over time, a phenomenon known as simplicity bias. Despite being widely observed across architectures, existing theoretical treatments lack a unifying framework. We present a theoretical framework that explains a simplicity bias arising from saddle-to-saddle learning dynamics for a general class of neural networks, incorporating fully-connected, convolutional, and attention-based architectures. Here, simple means expressible with few hidden units, i.e., hidden neurons, convolutional kernels, or attention heads. Specifically, we show that linear networks learn solutions of increasing rank, ReLU networks learn solutions with an increasing number of kinks, convolutional networks learn solutions with an increasing number of convolutional kernels, and self-attention models learn solutions with an increasing number of attention heads. By analyzing fixed points, invariant manifolds, and dynamics of gradient descent learning, we show that saddle-to-saddle dynamics operates by iteratively evolving near an invariant manifold, approaching a saddle, and switching to another invariant manifold. Our analysis also disentangles data-induced and initialization-induced saddle-to-saddle dynamics. In particular, the former leads to low-rank weights while the latter to sparse weights. Equipped with the theory, we predict the effects of data distribution and weight initialization on the duration and number of plateaus in learning. Overall, our theory offers a framework for understanding when and why gradient descent progressively learns increasingly complex solutions.

📄 PDF Abstract BibTeX arXiv:2512.20607

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Type-II Saddles and Probabilistic Stability of Stochastic Gradient Descent

2023-03-23 · Liu Ziyin, Botao Li, Tomer Galanti, Masahito Ueda

Characterizing and understanding the dynamics of stochastic gradient descent (SGD) around saddle points remains an open problem. We first show that saddle points in neural networks can be divided into two types, among wh…

Learning Theory

Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

2021-06-30 · Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler 외

The dynamics of Deep Linear Networks (DLNs) is dramatically affected by the variance $\sigma^2$ of the parameters at initialization $\theta_0$. For DLNs of width $w$, we show a phase transition w.r.t. the scaling $\gamma…

L2 Regularization

Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks

2026-04-07 · Guillaume Corlouer, Avi Semler, Alexander Strang, Alexander Gietelink Oldenziel arxiv

Deep linear networks (DLNs) are used as an analytically tractable model of the training dynamics of deep neural networks. While gradient descent in DLNs is known to exhibit saddle-to-saddle dynamics, the impact of stocha…

To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters

2026-02-28 · Sara Dragutinović, Yedi Zhang, Rajesh Ranganath arxiv

While Adam has long been the ubiquitous default optimizer for deep neural networks, Muon has recently seen rapid adoption due to its superior training speed. Although much of the literature focuses on validating the bene…

Saddle-to-Saddle Dynamics in Diagonal Linear Networks

2023-04-02 · NeurIPS 2023 11

In this paper we fully describe the trajectory of gradient flow over diagonal linear networks in the limit of vanishing initialisation. We show that the limiting flow successively jumps from a saddle of the training loss…

ARCIncremental Learning