paper-with-me

홈 › Papers

A Linearly-Convergent Stochastic L-BFGS Algorithm

2015-08-09 · Philipp Moritz, Robert Nishihara, Michael. I. Jordan

We propose a new stochastic L-BFGS algorithm and prove a linear convergence rate for strongly convex and smooth functions. Our algorithm draws heavily from a recent stochastic variant of L-BFGS proposed in Byrd et al. (2014) as well as a recent approach to variance reduction for stochastic gradient descent from Johnson and Zhang (2013). We demonstrate experimentally that our algorithm performs well on large-scale convex and non-convex optimization problems, exhibiting linear convergence and rapidly solving the optimization problems to high levels of precision. Furthermore, we show that our algorithm performs well for a wide-range of step sizes, often differing by several orders of magnitude.

📄 PDF Abstract BibTeX arXiv:1508.02087

Code (1)

crastogi/sLBFGS

Similar Papers 제목 키워드 기반

SPIRAL: A superlinearly convergent incremental proximal algorithm for nonconvex finite sum minimization

2022-07-17 · Pourya Behmandpoor, Puya Latafat, Andreas Themelis, Marc Moonen 외

We introduce SPIRAL, a SuPerlinearly convergent Incremental pRoximal ALgorithm, for solving nonconvex regularized finite sum problems under a relative smoothness assumption. Each iteration of SPIRAL consists of an inner …

Stochastic Steffensen method

2022-11-28 · Minda Zhao, Zehua Lai, Lek-Heng Lim

Is it possible for a first-order method, i.e., only first derivatives allowed, to be quadratically convergent? For univariate loss functions, the answer is yes -- the Steffensen method avoids second derivatives and is st…

Stochastic Optimization

Fast and Furious Convergence: Stochastic Second Order Methods under Interpolation

2019-10-11 · Si Yi Meng, Sharan Vaswani, Issam Laradji, Mark Schmidt 외

We consider stochastic second-order methods for minimizing smooth and strongly-convex functions under an interpolation condition satisfied by over-parameterized models. Under this condition, we show that the regularized …

Binary ClassificationSecond-order methods

mL-BFGS: A Momentum-based L-BFGS for Distributed Large-Scale Neural Network Optimization

2023-07-25 · Yue Niu, Zalan Fabian, Sunwoo Lee, Mahdi Soltanolkotabi 외

Quasi-Newton methods still face significant challenges in training large-scale neural networks due to additional compute costs in the Hessian related computations and instability issues in stochastic training. A well-kno…

Stochastic Optimization

Stochastic Damped L-BFGS with Controlled Norm of the Hessian Approximation

2020-12-10 · Sanae Lotfi, Tiphaine Bonniot de Ruisselet, Dominique Orban, Andrea Lodi

We propose a new stochastic variance-reduced damped L-BFGS algorithm, where we leverage estimates of bounds on the largest and smallest eigenvalues of the Hessian approximation to balance its quality and conditioning. Ou…

regression