paper-with-me

홈 › Papers

Sequential convergence of AdaGrad algorithm for smooth convex optimization

2020-11-24 · Cheik Traoré, Edouard Pauwels

We prove that the iterates produced by, either the scalar step size variant, or the coordinatewise variant of AdaGrad algorithm, are convergent sequences when applied to convex objective functions with Lipschitz gradient. The key insight is to remark that such AdaGrad sequences satisfy a variable metric quasi-Fej\'er monotonicity property, which allows to prove convergence.

📄 PDF Abstract BibTeX arXiv:2011.12341

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

AdaGrad AdaGrad is a stochastic optimization method that adapts the learning rate to the parameters. It performs smaller updates for parameters associated with frequently occurring…

Similar Papers 제목 키워드 기반

On the Convergence of AdaGrad(Norm) on $\R^{d}$: Beyond Convexity, Non-Asymptotic Rate and Acceleration

2022-09-29 · Zijian Liu, Ta Duy Nguyen, Alina Ene, Huy L. Nguyen

Existing analysis of AdaGrad and other adaptive methods for smooth convex optimization is typically for functions with bounded domain diameter. In unconstrained problems, previous works guarantee an asymptotic convergenc…

AdaGrad stepsizes: Sharp convergence over nonconvex landscapes

2018-06-05 · Rachel Ward, Xiaoxia Wu, Leon Bottou

Adaptive gradient methods such as AdaGrad and its variants update the stepsize in stochastic gradient descent on the fly according to the gradients received along the way; such methods have gained widespread use in large…

Stochastic Optimization

High Probability Bounds for a Class of Nonconvex Algorithms with AdaGrad Stepsize

2022-04-06 · ICLR 2022 4 · Ali Kavis, Kfir Yehuda Levy, Volkan Cevher

In this paper, we propose a new, simplified high probability analysis of AdaGrad for smooth, non-convex problems. More specifically, we focus on a particular accelerated gradient (AGD) template (Lan, 2020), through which…

Convergence of AdaGrad for Non-convex Objectives: Simple Proofs and Relaxed Assumptions

2023-05-29 · Bohan Wang, Huishuai Zhang, Zhi-Ming Ma, Wei Chen

We provide a simple convergence proof for AdaGrad optimizing non-convex objectives under only affine noise variance and bounded smoothness assumptions. The proof is essentially based on a novel auxiliary function $\xi$ t…

High Probability Convergence of Stochastic Gradient Methods

2023-02-28 · Zijian Liu, Ta Duy Nguyen, Thien Hang Nguyen, Alina Ene 외

In this work, we describe a generic approach to show convergence with high probability for both stochastic convex and non-convex optimization with sub-Gaussian noise. In previous works for convex optimization, either the…

Vocal Bursts Intensity Prediction