RMSprop can converge with proper hyper-parameter
Despite the existence of divergence examples, RMSprop remains one of the most popular algorithms in machine learning. Towards closing the gap between theory and practice, we prove that RMSprop can converge with proper choice of hyper-parameters under certain conditions. More specifically, we prove that when the hyper-parameter $\beta_2$ is large enough, the random shuffling version of RMSprop converges to a bounded region in general, and converges to a stationary point in the interpolation regime. It is worth mentioning that our results do not depend on "bounded gradient" assumption, which is often the key assumption utilized by existing theoretical work for RMSprop. Removing this assumption allows us to establish a phase transition from divergence to non-divergence for RMSProp. Finally, based on our theory, we conjecture that there is a critical threshold in practice, such that RMSprop generates reasonably good results only if $\beta_2\ge {\sf {th}}$. We provide empirical evidence about such a phase transition in our numerical experiments.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance
This paper provides the first tight convergence analyses for RMSProp and Adam in non-convex optimization under the most relaxed assumptions of coordinate-wise generalized smoothness and affine noise variance. We first an…
LEMMAConvergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
RMSProp and ADAM continue to be extremely popular algorithms for training neural nets but their theoretical convergence properties have remained unclear. Further, recent work has seemed to suggest that these algorithms h…
Convergence rates for the RMSprop optimizer with full control of the hyperparameters
Popular adaptive stochastic gradient descent (SGD) methods to train artificial intelligence (AI) systems include the RMSprop, the Adam, and the AdamW optimizers, where the adaptivity parts in Adam and AdamW basically jus…
Stochastic OptimizationConvergence of Steepest Descent and Adam under Non-Uniform Smoothness
Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine learning tasks. We generalize this assumption to objectives whose c…
Reinforcement LearningExploring the Optimized Value of Each Hyperparameter in Various Gradient Descent Algorithms
In the recent years, various gradient descent algorithms including the methods of gradient descent, gradient descent with momentum, adaptive gradient (AdaGrad), root-mean-square propagation (RMSProp) and adaptive moment …
Model Optimization