paper-with-me

Papers

Nesterov acceleration despite very noisy gradients

2023-02-10 · Kanan Gupta, Jonathan W. Siegel, Stephan Wojtowytsch

We present a generalization of Nesterov's accelerated gradient descent algorithm. Our algorithm (AGNES) provably achieves acceleration for smooth convex and strongly convex minimization tasks with noisy gradient estimates if the noise intensity is proportional to the magnitude of the gradient at every point. Nesterov's method converges at an accelerated rate if the constant of proportionality is below 1, while AGNES accommodates any signal-to-noise ratio. The noise model is motivated by applications in overparametrized machine learning. AGNES requires only two parameters in convex and three in strongly convex minimization tasks, improving on existing methods. We further provide clear geometric interpretations and heuristics for the choice of parameters.

📄 PDF Abstract BibTeX arXiv:2302.05515

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Continuized View on Nesterov Acceleration for Stochastic Gradient Descent and Randomized Gossip

2021-06-10 · Mathieu Even, Raphaël Berthier, Francis Bach, Nicolas Flammarion 외

We introduce the continuized Nesterov acceleration, a close variant of Nesterov acceleration whose variables are indexed by a continuous time parameter. The two variables continuously mix following a linear ordinary diff…

Continuized Accelerations of Deterministic and Stochastic Gradient Descents, and of Gossip Algorithms

2021-12-01 · NeurIPS 2021 12 · Mathieu Even, Raphaël Berthier, Francis Bach, Nicolas Flammarion 외

We introduce the ``continuized'' Nesterov acceleration, a close variant of Nesterov acceleration whose variables are indexed by a continuous time parameter. The two variables continuously mix following a linear ordinary …

EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization

2026-05-25 · Chung-Yiu Yau, Dawei Li, Athanasios Glentis, Valentyn Boreiko 외 arxiv

Lookahead-based acceleration methods, such as Nesterov's momentum, are widely used in optimization, but they often become unreliable in deep learning training mainly due to stochastic gradient noise and non-convex loss l…

Accelerated Reinforcement Learning

2017-10-23 · K. Lakshmanan

Policy gradient methods are widely used in reinforcement learning algorithms to search for better policies in the parameterized policy space. They do gradient search in the policy space and are known to converge very slo…

Policy Gradient Methodsreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2

On the Convergence of Nesterov's Accelerated Gradient Method in Stochastic Settings

2020-02-27 · ICML 2020 1 · Mahmoud Assran, Michael Rabbat

We study Nesterov's accelerated gradient method with constant step-size and momentum parameters in the stochastic approximation setting (unbiased gradients with bounded variance) and the finite-sum setting (where randomn…