paper-with-me

Papers

The Power of Factorial Powers: New Parameter settings for (Stochastic) Optimization

2020-06-01 · Aaron Defazio, Robert M. Gower

The convergence rates for convex and non-convex optimization methods depend on the choice of a host of constants, including step sizes, Lyapunov function constants and momentum constants. In this work we propose the use of factorial powers as a flexible tool for defining constants that appear in convergence proofs. We list a number of remarkable properties that these sequences enjoy, and show how they can be applied to convergence proofs to simplify or improve the convergence rates of the momentum method, accelerated gradient and the stochastic variance reduced method (SVRG).

📄 PDF Abstract BibTeX arXiv:2006.01244

Code (0)

등록된 구현이 없습니다.

Tasks

Stochastic Optimization

Similar Papers 제목 키워드 기반

PowerSGD: Powered Stochastic Gradient Descent Methods for Accelerated Non-Convex Optimization

2019-09-25 · Jun Liu, Beitong Zhou, Weigao Sun, Ruijuan Chen 외

In this paper, we propose a novel technique for improving the stochastic gradient descent (SGD) method to train deep networks, which we term \emph{PowerSGD}. The proposed PowerSGD method simply raises the stochastic grad…

From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees

2025-09-14 · Shengping Xie, Chuyan Chen, Kun Yuan arxiv

Low-rank gradient compression methods, such as PowerSGD, have gained attention in communication-efficient distributed optimization. However, the convergence guarantees of PowerSGD remain unclear, particularly in stochast…

Distributed Optimization

PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent

2026-05-11 · Yao Lu, Dengdong Fan, Shixun Zhang, Yonghong Tian arxiv

Adaptive optimizers, most notably Adam, have become the default standard for training large-scale neural networks such as Transformers. These methods maintain running estimates of gradient first and second moments, incur…

Stochastic Optimization

Scaling Factorial Hidden Markov Models: Stochastic Variational Inference without Messages

2016-08-12 · NeurIPS 2016 12 · Yin Cheng Ng, Pawel Chilinski, Ricardo Silva

Factorial Hidden Markov Models (FHMMs) are powerful models for sequential data but they do not scale well with long sequences. We propose a scalable inference and learning algorithm for FHMMs that draws on ideas from the…

Variational Inference

Stochastic Analysis of the Diffusion Least Mean Square and Normalized Least Mean Square Algorithms for Cyclostationary White Gaussian and Non-Gaussian Inputs

2021-08-05 · Eweda Eweda, Neil J. Bershad, Jose C. M. Bermudez

The diffusion least mean square (DLMS) and the diffusion normalized least mean square (DNLMS) algorithms are analyzed for a network having a fusion center. This structure reduces the dimensionality of the resulting stoch…