paper-with-me

홈 › Papers

Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability

2026-04-03 · Eric Gan arxiv

Empirically, modern deep learning training often occurs at the Edge of Stability (EoS), where the sharpness of the loss exceeds the threshold below which classical convergence analysis applies. Despite recent progress, existing theoretical explanations of EoS either rely on restrictive assumptions or focus on specific squared-loss-type objectives. In this work, we introduce and study a structural property of loss functions that we term product-stability. We show that for losses with product-stable minima, gradient descent applied to objectives of the form $(x,y) \mapsto l(xy)$ can provably converge to the local minimum even when training in the EoS regime. This framework substantially generalizes prior results and applies to a broad class of losses, including binary cross entropy. Using bifurcation diagrams, we characterize the resulting training dynamics, explain the emergence of stable oscillations, and precisely quantify the sharpness at convergence. Together, our results offer a principled explanation for stable EoS training for a wider class of loss functions.

📄 PDF Abstract BibTeX arXiv:2604.02653

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent

2023-11-28 · Frederik Köhne, Leonie Kreis, Anton Schiela, Roland Herzog

This paper proposes a novel approach to adaptive step sizes in stochastic gradient descent (SGD) by utilizing quantities that we have identified as numerically traceable -- the Lipschitz constant for gradients and a conc…

image-classificationImage ClassificationStochastic Optimization

Federated Accelerated Stochastic Gradient Descent

2020-06-16 · NeurIPS 2020 12 · Honglin Yuan, Tengyu Ma

We propose Federated Accelerated Stochastic Gradient Descent (FedAc), a principled acceleration of Federated Averaging (FedAvg, also known as Local SGD) for distributed optimization. FedAc is the first provable accelerat…

Distributed Optimization

Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent

2025-02-16 · Ya-Chi Chu, Wenzhi Gao, Yinyu Ye, Madeleine Udell

This paper investigates the convergence properties of the hypergradient descent method (HDM), a 25-year-old heuristic originally proposed for adaptive stepsize selection in stochastic first-order methods. We provide the …

A Bifurcation Theory Framework for Gradient Descent on the Edge of Stability

2026-06-14 · Eric Gan arxiv

The Edge of Stability (EoS) phenomenon, where gradient descent operates with sharpness exceeding the classical convergence threshold yet the loss decreases over long timescales, is ubiquitous in modern deep learning but …

Convergence Rates for Gradient Descent on the Edge of Stability in Overparametrised Least Squares

2025-10-20 · Lachlan Ewen MacDonald, Hancheng Min, Leandro Palma, Salma Tarmoun 외 arxiv

Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or ``stable", regime. In contrast, gradient descent on neural networks is frequently per…