paper-with-me

홈 › Papers

RMSprop can converge with proper hyper-parameter

2021-01-01 · ICLR 2021 1 · Naichen Shi, Dawei Li, Mingyi Hong, Ruoyu Sun

Despite the existence of divergence examples, RMSprop remains one of the most popular algorithms in machine learning. Towards closing the gap between theory and practice, we prove that RMSprop can converge with proper choice of hyper-parameters under certain conditions. More specifically, we prove that when the hyper-parameter $\beta_2$ is large enough, the random shuffling version of RMSprop converges to a bounded region in general, and converges to a stationary point in the interpolation regime. It is worth mentioning that our results do not depend on "bounded gradient" assumption, which is often the key assumption utilized by existing theoretical work for RMSprop. Removing this assumption allows us to establish a phase transition from divergence to non-divergence for RMSProp. Finally, based on our theory, we conjecture that there is a critical threshold in practice, such that RMSprop generates reasonably good results only if $\beta_2\ge {\sf {th}}$. We provide empirical evidence about such a phase transition in our numerical experiments.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

RMSProp RMSProp is an unpublished adaptive learning rate optimizer proposed by Geoff Hinton. The motivation…

Similar Papers 제목 키워드 기반

Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance

2024-04-01 · Qi Zhang, Yi Zhou, Shaofeng Zou

This paper provides the first tight convergence analyses for RMSProp and Adam in non-convex optimization under the most relaxed assumptions of coordinate-wise generalized smoothness and affine noise variance. We first an…

LEMMA

Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration

2018-07-18 · ICLR 2019 5 · Soham De, Anirbit Mukherjee, Enayat Ullah

RMSProp and ADAM continue to be extremely popular algorithms for training neural nets but their theoretical convergence properties have remained unclear. Further, recent work has seemed to suggest that these algorithms h…

Convergence rates for the RMSprop optimizer with full control of the hyperparameters

2026-08-31 · Steffen Dereich, Arnulf Jentzen arxiv

Popular adaptive stochastic gradient descent (SGD) methods to train artificial intelligence (AI) systems include the RMSprop, the Adam, and the AdamW optimizers, where the adaptivity parts in Adam and AdamW basically jus…

Stochastic Optimization

Convergence of Steepest Descent and Adam under Non-Uniform Smoothness

2026-05-28 · Sharan Vaswani, Yifan Sun, Reza Babanezhad arxiv

Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine learning tasks. We generalize this assumption to objectives whose c…

Reinforcement Learning

Exploring the Optimized Value of Each Hyperparameter in Various Gradient Descent Algorithms

2022-12-23 · Abel C. H. Chen

In the recent years, various gradient descent algorithms including the methods of gradient descent, gradient descent with momentum, adaptive gradient (AdaGrad), root-mean-square propagation (RMSProp) and adaptive moment …

Model Optimization