paper-with-me

홈 › Papers

Nesterov acceleration in benignly non-convex landscapes

2024-10-10 · Kanan Gupta, Stephan Wojtowytsch

While momentum-based optimization algorithms are commonly used in the notoriously non-convex optimization problems of deep learning, their analysis has historically been restricted to the convex and strongly convex setting. In this article, we partially close this gap between theory and practice and demonstrate that virtually identical guarantees can be obtained in optimization problems with a `benign' non-convexity. We show that these weaker geometric assumptions are well justified in overparametrized deep learning, at least locally. Variations of this result are obtained for a continuous time model of Nesterov's accelerated gradient descent algorithm (NAG), the classical discrete time version of NAG, and versions of NAG with stochastic gradient estimates with purely additive noise and with noise that exhibits both additive and multiplicative scaling.

📄 PDF Abstract BibTeX arXiv:2410.08395

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Learning

Similar Papers 제목 키워드 기반

EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization

2026-05-25 · Chung-Yiu Yau, Dawei Li, Athanasios Glentis, Valentyn Boreiko 외 arxiv

Lookahead-based acceleration methods, such as Nesterov's momentum, are widely used in optimization, but they often become unreliable in deep learning training mainly due to stochastic gradient noise and non-convex loss l…

Analytical Study of Momentum-Based Acceleration Methods in Paradigmatic High-Dimensional Non-Convex Problems

2021-02-23 · NeurIPS 2021 12 · Stefano Sarao Mannelli, Pierfrancesco Urbani

The optimization step in many machine learning problems rarely relies on vanilla gradient descent but it is common practice to use momentum-based accelerated methods. Despite these algorithms being widely applied to arbi…

Numerical Integration

Nesterov Acceleration of Alternating Least Squares for Canonical Tensor Decomposition: Momentum Step Size Selection and Restart Mechanisms

2018-10-13 · Drew Mitchell, Nan Ye, Hans De Sterck

We present Nesterov-type acceleration techniques for Alternating Least Squares (ALS) methods applied to canonical tensor decomposition. While Nesterov acceleration turns gradient descent into an optimal first-order metho…

Tensor Decomposition

Accelerated Reinforcement Learning

2017-10-23 · K. Lakshmanan

Policy gradient methods are widely used in reinforcement learning algorithms to search for better policies in the parameterized policy space. They do gradient search in the policy space and are known to converge very slo…

Policy Gradient Methodsreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2

Nesterov acceleration despite very noisy gradients

2023-02-10 · Kanan Gupta, Jonathan W. Siegel, Stephan Wojtowytsch

We present a generalization of Nesterov's accelerated gradient descent algorithm. Our algorithm (AGNES) provably achieves acceleration for smooth convex and strongly convex minimization tasks with noisy gradient estimate…