paper-with-me

홈 › Papers

On the Convergence Proof of AMSGrad and a New Version

2019-04-07 · Tran Thi Phuong, Le Trieu Phong

The adaptive moment estimation algorithm Adam (Kingma and Ba) is a popular optimizer in the training of deep neural networks. However, Reddi et al. have recently shown that the convergence proof of Adam is problematic and proposed a variant of Adam called AMSGrad as a fix. In this paper, we show that the convergence proof of AMSGrad is also problematic. Concretely, the problem in the convergence proof of AMSGrad is in handling the hyper-parameters, treating them as equal while they are not. This is also the neglected issue in the convergence proof of Adam. We provide an explicit counter-example of a simple convex optimization setting to show this neglected issue. Depending on manipulating the hyper-parameters, we present various fixes for this issue. We provide a new convergence proof for AMSGrad as the first fix. We also propose a new version of AMSGrad called AdamX as another fix. Our experiments on the benchmark dataset also support our theoretical results.

📄 PDF Abstract BibTeX arXiv:1904.03590

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Adam 설명 없음

Similar Papers 제목 키워드 기반

Towards Deep Robot Learning with Optimizer applicable to Non-stationary Problems

2020-07-31 · Taisuke Kobayashi

This paper proposes a new optimizer for deep learning, named d-AmsGrad. In the real-world data, noise and outliers cannot be excluded from dataset to be used for learning robot skills. This problem is especially striking…

MixML: A Unified Analysis of Weakly Consistent Parallel Learning

2020-05-14 · Yucheng Lu, Jack Nash, Christopher De Sa

Parallelism is a ubiquitous method for accelerating machine learning algorithms. However, theoretical analysis of parallel learning is usually done in an algorithm- and protocol-specific setting, giving little insight ab…

BIG-bench Machine Learning

A Theoretical and Experimental Study of a Novel Adaptive Learning Algorithm

2026-05-28 · Sakshi Kumari, Shyam Kumar M, Sushmitha P arxiv

A crucial component of machine learning algorithms is minimizing loss functions with less computational cost and less oscillations. While adaptive learning rate-based optimizers have been widely used for real-world tasks…

Analysis of Q-learning with Adaptation and Momentum Restart for Gradient Descent

2020-07-15 · Bowen Weng, Huaqing Xiong, Yingbin Liang, Wei zhang

Existing convergence analyses of Q-learning mostly focus on the vanilla stochastic gradient descent (SGD) type of updates. Despite the Adaptive Moment Estimation (Adam) has been commonly used for practical Q-learning alg…

Atari GamesQ-Learning

An Optimistic Acceleration of AMSGrad for Nonconvex Optimization

2019-03-04 · ICLR 2020 1 · Jun-Kun Wang, Xiaoyun Li, Belhal Karimi, Ping Li

We propose a new variant of AMSGrad, a popular adaptive gradient based optimization algorithm widely used for training deep neural networks. Our algorithm adds prior knowledge about the sequence of consecutive mini-batch…