paper-with-me

홈 › Papers

Benign Oscillation of Stochastic Gradient Descent with Large Learning Rates

2023-10-26 · Miao Lu, Beining Wu, Xiaodong Yang, Difan Zou

In this work, we theoretically investigate the generalization properties of neural networks (NN) trained by stochastic gradient descent (SGD) algorithm with large learning rates. Under such a training regime, our finding is that, the oscillation of the NN weights caused by the large learning rate SGD training turns out to be beneficial to the generalization of the NN, which potentially improves over the same NN trained by SGD with small learning rates that converges more smoothly. In view of this finding, we call such a phenomenon "benign oscillation". Our theory towards demystifying such a phenomenon builds upon the feature learning perspective of deep learning. Specifically, we consider a feature-noise data generation model that consists of (i) weak features which have a small $\ell_2$-norm and appear in each data point; (ii) strong features which have a larger $\ell_2$-norm but only appear in a certain fraction of all data points; and (iii) noise. We prove that NNs trained by oscillating SGD with a large learning rate can effectively learn the weak features in the presence of those strong features. In contrast, NNs trained by SGD with a small learning rate can only learn the strong features but makes little progress in learning the weak features. Consequently, when it comes to the new testing data which consist of only weak features, the NN trained by oscillating SGD with a large learning rate could still make correct predictions consistently, while the NN trained by small learning rate SGD fails. Our theory sheds light on how large learning rate training benefits the generalization of NNs. Experimental results demonstrate our finding on "benign oscillation".

📄 PDF Abstract BibTeX arXiv:2310.17074

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System

2026-07-06 · Thomas Hofmann arxiv

Many phenomena of deep learning are dynamical: they concern not only which minima exist, but how gradient descent reaches, avoids, or selects among them. Edge-of-stability behavior, sharpness oscillations, catapult phase…

SGD at the Edge of Stability: Stochastic Stabilization with Large Learning Rates

2026-06-29 · Konstantinos Emmanouilidis, Lachlan MacDonald, Salma Tarmoun, Rene Vidal arxiv

Modern deep learning has been shown to operate at the edge of stability, routinely using learning rates far larger than those justified by classical optimization theory. Most prior analyses of the edge of stability pheno…

Complex Stochastic Gradient Descent and Directional Bias in Reproducing Kernel Hilbert Spaces

2026-04-24 · Natanael Alpay, Emeric Battaglia arxiv

Stochastic Gradient Descent (SGD) is a known stochastic iterative method popular for large-scale convex optimization problems due to its simple implementation and scalability. Some objectives, such as those found in comp…

Tackling benign nonconvexity with smoothing and stochastic gradients

2022-02-18 · Harsh Vardhan, Sebastian U. Stich

Non-convex optimization problems are ubiquitous in machine learning, especially in Deep Learning. While such complex problems can often be successfully optimized in practice by using stochastic gradient descent (SGD), th…

More Optimal Fractional-Order Stochastic Gradient Descent for Non-Convex Optimization Problems

2025-05-05 · Mohammad Partohaghighi, Roummel Marcia, YangQuan Chen

Fractional-order stochastic gradient descent (FOSGD) leverages fractional exponents to capture long-memory effects in optimization. However, its utility is often limited by the difficulty of tuning and stabilizing these …