paper-with-me

홈 › Papers

Convergence of optimizers implies eigenvalues filtering at equilibrium

2025-10-10 · Jerome Bolte, Quoc-Tung Le, Edouard Pauwels arxiv

Ample empirical evidence in deep neural network training suggests that a variety of optimizers tend to find nearly global optima. In this article, we adopt the reversed perspective that convergence to an arbitrary point is assumed rather than proven, focusing on the consequences of this assumption. From this viewpoint, in line with recent advances on the edge-of-stability phenomenon, we argue that different optimizers effectively act as eigenvalue filters determined by their hyperparameters. Specifically, the standard gradient descent method inherently avoids the sharpest minima, whereas Sharpness-Aware Minimization (SAM) algorithms go even further by actively favoring wider basins. Inspired by these insights, we propose two novel algorithms that exhibit enhanced eigenvalue filtering, effectively promoting wider minima. Our theoretical analysis leverages a generalized Hadamard--Perron stable manifold theorem and applies to general semialgebraic $C^2$ functions, without requiring additional non-degeneracy conditions or global Lipschitz bound assumptions. We support our conclusions with numerical experiments on feed-forward neural networks.

📄 PDF Abstract BibTeX arXiv:2510.09034

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Number of Steps Needed for Nonconvex Optimization of a Deep Learning Optimizer is a Rational Function of Batch Size

2021-08-26 · Hideaki Iiduka

Recently, convergence as well as convergence rate analyses of deep learning optimizers for nonconvex optimization have been widely studied. Meanwhile, numerical evaluations for the optimizers have precisely clarified the…

Flowers of immortality

2022-10-24 · Thomas Fink, Yang-Hui He

There has been a recent surge of interest in what causes aging. This has been matched by unprecedented research investment in the field from tech companies. But, despite considerable effort from a broad range of research…

Augmented Synchronization of Power Systems

2021-06-24 · Peng Yang, Feng Liu, Tao Liu, David J. Hill

Power system transient stability has been translated into a Lyapunov stability problem of the post-disturbance equilibrium for decades. Despite substantial results, conventional theories suffer from the stringent require…

Domain Adversarial Training: A Game Perspective

2022-02-10 · ICLR 2022 4 · David Acuna, Marc T Law, Guojun Zhang, Sanja Fidler

The dominant line of work in domain adaptation has focused on learning invariant representations using domain-adversarial training. In this paper, we interpret this approach from a game theoretical perspective. Defining …

Domain Adaptation

STay-ON-the-Ridge: Guaranteed Convergence to Local Minimax Equilibrium in Nonconvex-Nonconcave Games

2022-10-18 · Constantinos Daskalakis, Noah Golowich, Stratis Skoulakis, Manolis Zampetakis

Min-max optimization problems involving nonconvex-nonconcave objectives have found important applications in adversarial training and other multi-agent learning settings. Yet, no known gradient descent-based method is gu…