paper-with-me

홈 › Papers

Saddle-to-Saddle Dynamics in Diagonal Linear Networks

2023-04-02 · NeurIPS 2023 11

In this paper we fully describe the trajectory of gradient flow over diagonal linear networks in the limit of vanishing initialisation. We show that the limiting flow successively jumps from a saddle of the training loss to another until reaching the minimum $\ell_1$-norm solution. This saddle-to-saddle dynamics translates to an incremental learning process as each saddle corresponds to the minimiser of the loss constrained to an active set outside of which the coordinates must be zero. We explicitly characterise the visited saddles as well as the jumping times through a recursive algorithm reminiscent of the LARS algorithm used for computing the Lasso path. Our proof leverages a convenient arc-length time-reparametrisation which enables to keep track of the heteroclinic transitions between the jumps. Our analysis requires negligible assumptions on the data, applies to both under and overparametrised settings and covers complex cases where there is no monotonicity of the number of active coordinates. We provide numerical experiments to support our findings.

📄 PDF Abstract BibTeX arXiv:2304.00488

Code (0)

등록된 구현이 없습니다.

Tasks

ARCIncremental Learning

Methods 이 논문이 사용한 방법론

LARS Layer-wise Adaptive Rate Scaling, or LARS, is a large batch optimization technique. There are two notable differences between LARS and other adaptive algorithms such as…

Similar Papers 제목 키워드 기반

Never Saddle for Reparameterized Steepest Descent as Mirror Flow

2026-03-02 · Tom Jacobs, Chao Zhou, Rebekka Burkholz arxiv

How does the choice of optimization algorithm shape a model's ability to learn features? To address this question for steepest descent methods --including sign descent, which is closely related to Adam --we introduce ste…

Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

2021-06-30 · Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler 외

The dynamics of Deep Linear Networks (DLNs) is dramatically affected by the variance $\sigma^2$ of the parameters at initialization $\theta_0$. For DLNs of width $w$, we show a phase transition w.r.t. the scaling $\gamma…

L2 Regularization

Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks

2026-04-07 · Guillaume Corlouer, Avi Semler, Alexander Strang, Alexander Gietelink Oldenziel arxiv

Deep linear networks (DLNs) are used as an analytically tractable model of the training dynamics of deep neural networks. While gradient descent in DLNs is known to exhibit saddle-to-saddle dynamics, the impact of stocha…

Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures

2025-12-23 · Yedi Zhang, Andrew Saxe, Peter E. Latham arxiv

Neural networks trained with gradient descent often learn solutions of increasing complexity over time, a phenomenon known as simplicity bias. Despite being widely observed across architectures, existing theoretical trea…

Type-II Saddles and Probabilistic Stability of Stochastic Gradient Descent

2023-03-23 · Liu Ziyin, Botao Li, Tomer Galanti, Masahito Ueda

Characterizing and understanding the dynamics of stochastic gradient descent (SGD) around saddle points remains an open problem. We first show that saddle points in neural networks can be divided into two types, among wh…

Learning Theory