paper-with-me

홈 › Papers

Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization

2021-08-25 · Difan Zou, Yuan Cao, Yuanzhi Li, Quanquan Gu

Adaptive gradient methods such as Adam have gained increasing popularity in deep learning optimization. However, it has been observed that compared with (stochastic) gradient descent, Adam can converge to a different solution with a significantly worse test error in many deep learning applications such as image classification, even with a fine-tuned regularization. In this paper, we provide a theoretical explanation for this phenomenon: we show that in the nonconvex setting of learning over-parameterized two-layer convolutional neural networks starting from the same random initialization, for a class of data distributions (inspired from image data), Adam and gradient descent (GD) can converge to different global solutions of the training objective with provably different generalization errors, even with weight decay regularization. In contrast, we show that if the training objective is convex, and the weight decay regularization is employed, any optimization algorithms including Adam and GD will converge to the same solution if the training is successful. This suggests that the inferior generalization performance of Adam is fundamentally tied to the nonconvex landscape of deep learning optimization.

📄 PDF Abstract BibTeX arXiv:2108.11371

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Learningimage-classificationImage Classification

Methods 이 논문이 사용한 방법론

Weight Decay 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

Implicit Bias of AdamW: $\ell_\infty$ Norm Constrained Optimization

2024-04-05 · Shuo Xie, Zhiyuan Li

Adam with decoupled weight decay, also known as AdamW, is widely acclaimed for its superior performance in language modeling tasks, surpassing Adam with $\ell_2$ regularization in terms of generalization and optimization…

Language ModelingLanguage Modelling

Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks

2025-10-13 · Xuan Tang, Han Zhang, Yuan Cao, Difan Zou arxiv

Adam is a popular and widely used adaptive gradient method in deep learning, which has also received tremendous focus in theoretical research. However, most existing theoretical work primarily analyzes its full-batch ver…

Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration

2018-07-18 · ICLR 2019 5 · Soham De, Anirbit Mukherjee, Enayat Ullah

RMSProp and ADAM continue to be extremely popular algorithms for training neural nets but their theoretical convergence properties have remained unclear. Further, recent work has seemed to suggest that these algorithms h…

Decoupled Weight Decay Regularization

2017-11-14 · ICLR 2019 5 · Ilya Loshchilov, Frank Hutter

L$_2$ regularization and weight decay regularization are equivalent for standard stochastic gradient descent (when rescaled by the learning rate), but as we demonstrate this is \emph{not} the case for adaptive gradient a…

image-classificationImage Classification

Understanding Dynamics of Adam in Zero-Sum Games: An ODE Approach

2026-05-19 · Yi Feng, Weiming Ou, Xiao Wang arxiv

The remarkable success of the Adam in training neural networks has naturally led to the widespread use of its descent-ascent counterpart, Adam-DA, for solving zero-sum games. Despite its popularity in practice, a rigorou…