paper-with-me

홈 › Papers

Divergence of the ADAM algorithm with fixed-stepsize: a (very) simple example

2023-08-01 · Ph. L. Toint

A very simple unidimensional function with Lipschitz continuous gradient is constructed such that the ADAM algorithm with constant stepsize, started from the origin, diverges when applied to minimize this function in the absence of noise on the gradient. Divergence occurs irrespective of the choice of the method parameters.

📄 PDF Abstract BibTeX arXiv:2308.00720

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Adam 설명 없음

Similar Papers 제목 키워드 기반

A Langevin sampling algorithm inspired by the Adam optimizer

2025-04-26 · Benedict Leimkuhler, René Lohmann, Peter Whalley

We present a framework for adaptive-stepsize MCMC sampling based on time-rescaled Langevin dynamics, in which the stepsize variation is dynamically driven by an additional degree of freedom. Our approach augments the pha…

Expectigrad: Fast Stochastic Optimization with Robust Convergence Properties

2020-10-03 · Brett Daley, Christopher Amato

Many popular adaptive gradient methods such as Adam and RMSProp rely on an exponential moving average (EMA) to normalize their stepsizes. While the EMA makes these methods highly responsive to new gradient information, r…

Stochastic Optimization

Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian Sampling

2020-02-15 · Huaqing Xiong, Tengyu Xu, Yingbin Liang, Wei zhang

Despite the wide applications of Adam in reinforcement learning (RL), the theoretical convergence of Adam-type RL algorithms has not been established. This paper provides the first such convergence analysis for two funda…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

Adaptive Stepsizing for Stochastic Gradient Langevin Dynamics in Bayesian Neural Networks

2025-11-11 · Rajit Rajpal, Benedict Leimkuhler, Yuanhao Jiang arxiv

Bayesian neural networks (BNNs) require scalable sampling algorithms to approximate posterior distributions over parameters. Existing stochastic gradient Markov Chain Monte Carlo (SGMCMC) methods are highly sensitive to …

Image Classification

Weak Convergence Properties of Constrained Emphatic Temporal-difference Learning with Constant and Slowly Diminishing Stepsize

2015-11-23 · Huizhen Yu

We consider the emphatic temporal-difference (TD) algorithm, ETD($\lambda$), for learning the value functions of stationary policies in a discounted, finite state and action Markov decision process. The ETD($\lambda$) al…