paper-with-me

홈 › Papers

Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails

2026-03-03 · Ruinan Jin, Yingbin Liang, Shaofeng Zou arxiv

Despite Adam demonstrating faster empirical convergence than SGD in many applications, much of the existing theory yields guarantees essentially comparable to those of SGD, leaving the empirical performance gap insufficiently explained. In this paper, we uncover a key second-moment normalization in Adam and develop a stopping-time/martingale analysis that provably distinguishes Adam from SGD under the classical bounded variance model (a second moment assumption). In particular, we establish the first theoretical separation between the high-probability convergence behaviors of the two methods: Adam achieves a $δ^{-1/2}$ dependence on the confidence parameter $δ$, whereas corresponding high-probability guarantee for SGD necessarily incurs at least a $δ^{-1}$ dependence.

📄 PDF Abstract BibTeX arXiv:2603.03099

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Adaptive Optimization via Momentum on Variance-Normalized Gradients

2026-02-10 · Francisco Patitucci, Aryan Mokhtari arxiv

We introduce MVN-Grad (Momentum on Variance-Normalized Gradients), an Adam-style optimizer that improves stability and performance by combining two complementary ideas: variance-based normalization and momentum applied a…

Image Classification

Training Deep Networks with Stochastic Gradient Normalized by Layerwise Adaptive Second Moments

2020-01-01 · ICLR 2020 1 · Boris Ginsburg, Patrice Castonguay, Oleksii Hrinchuk, Oleksii Kuchaiev 외

We propose NovoGrad, an adaptive stochastic gradient descent method with layer-wise gradient normalization and decoupled weight decay. In our experiments on neural networks for image classification, speech recognition, m…

General Classificationimage-classificationImage ClassificationLanguage Modeling+5

A Minimalist Optimizer Design for LLM Pretraining

2025-06-20 · Athanasios Glentis, Jiaxiang Li, Andi Han, Mingyi Hong

Training large language models (LLMs) typically relies on adaptive optimizers such as Adam, which require significant memory to maintain first- and second-moment matrices, known as optimizer states. While recent works su…

ADOPT: Modified Adam Can Converge with Any $β_2$ with the Optimal Rate

2024-11-05 · Shohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima 외

Adam is one of the most popular optimization algorithms in deep learning. However, it is known that Adam does not converge in theory unless choosing a hyperparameter, i.e., $\beta_2$, in a problem-dependent manner. There…

Deep Reinforcement Learningimage-classificationImage Classification

FedAdamW: A Communication-Efficient Optimizer with Convergence and Generalization Guarantees for Federated Large Models

2025-10-31 · Junkang Liu, Fanhua Shang, Hongying Liu, Yuxuan Tian 외 arxiv

AdamW has become one of the most effective optimizers for training large-scale models. We have also observed its effectiveness in the context of federated learning (FL). However, directly applying AdamW in federated lear…

Federated Learning