paper-with-me

홈 › Papers

High Probability Convergence of Adam Under Unbounded Gradients and Affine Variance Noise

2023-11-03 · Yusu Hong, Junhong Lin

In this paper, we study the convergence of the Adaptive Moment Estimation (Adam) algorithm under unconstrained non-convex smooth stochastic optimizations. Despite the widespread usage in machine learning areas, its theoretical properties remain limited. Prior researches primarily investigated Adam's convergence from an expectation view, often necessitating strong assumptions like uniformly stochastic bounded gradients or problem-dependent knowledge in prior. As a result, the applicability of these findings in practical real-world scenarios has been constrained. To overcome these limitations, we provide a deep analysis and show that Adam could converge to the stationary point in high probability with a rate of $\mathcal{O}\left({\rm poly}(\log T)/\sqrt{T}\right)$ under coordinate-wise "affine" variance noise, not requiring any bounded gradient assumption and any problem-dependent knowledge in prior to tune hyper-parameters. Additionally, it is revealed that Adam confines its gradients' magnitudes within an order of $\mathcal{O}\left({\rm poly}(\log T)\right)$. Finally, we also investigate a simplified version of Adam without one of the corrective terms and obtain a convergence rate that is adaptive to the noise level.

📄 PDF Abstract BibTeX arXiv:2311.02000

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Adam 설명 없음

Similar Papers 제목 키워드 기반

On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions

2024-02-06 · Yusu Hong, Junhong Lin

The Adaptive Momentum Estimation (Adam) algorithm is highly effective in training various deep learning tasks. Despite this, there's limited theoretical understanding for Adam, especially when focusing on its vanilla for…

Stochastic Optimization

Clipping Improves Adam-Norm and AdaGrad-Norm when the Noise Is Heavy-Tailed

2024-06-06 · Savelii Chezhegov, Yaroslav Klyukin, Andrei Semenov, Aleksandr Beznosikov 외

Methods with adaptive stepsizes, such as AdaGrad and Adam, are essential for training modern Deep Learning models, especially Large Language Models. Typically, the noise in the stochastic gradients is heavy-tailed for th…

Stochastic Optimization

On the Convergence of Adam-Type Algorithm for Bilevel Optimization under Unbounded Smoothness

2025-03-05 · Xiaochuan Gong, Jie Hao, Mingrui Liu

Adam has become one of the most popular optimizers for training modern deep neural networks, such as transformers. However, its applicability is largely restricted to single-level optimization problems. In this paper, we…

Bilevel OptimizationLEMMAMeta-Learning

Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails

2026-03-03 · Ruinan Jin, Yingbin Liang, Shaofeng Zou arxiv

Despite Adam demonstrating faster empirical convergence than SGD in many applications, much of the existing theory yields guarantees essentially comparable to those of SGD, leaving the empirical performance gap insuffici…

Robustness to Unbounded Smoothness of Generalized SignSGD

2022-08-23 · Michael Crawshaw, Mingrui Liu, Francesco Orabona, Wei zhang 외

Traditional analyses in non-convex optimization typically rely on the smoothness assumption, namely requiring the gradients to be Lipschitz. However, recent evidence shows that this smoothness condition does not capture …