paper-with-me

Papers

Fast and Furious Convergence: Stochastic Second Order Methods under Interpolation

2019-10-11 · Si Yi Meng, Sharan Vaswani, Issam Laradji, Mark Schmidt, Simon Lacoste-Julien

We consider stochastic second-order methods for minimizing smooth and strongly-convex functions under an interpolation condition satisfied by over-parameterized models. Under this condition, we show that the regularized subsampled Newton method (R-SSN) achieves global linear convergence with an adaptive step-size and a constant batch-size. By growing the batch size for both the subsampled gradient and Hessian, we show that R-SSN can converge at a quadratic rate in a local neighbourhood of the solution. We also show that R-SSN attains local linear convergence for the family of self-concordant functions. Furthermore, we analyze stochastic BFGS algorithms in the interpolation setting and prove their global linear convergence. We empirically evaluate stochastic L-BFGS and a "Hessian-free" implementation of R-SSN for binary classification on synthetic, linearly-separable datasets and real datasets under a kernel mapping. Our experimental results demonstrate the fast convergence of these methods, both in terms of the number of iterations and wall-clock time.

📄 PDF Abstract BibTeX arXiv:1910.04920

Code (1)

IssamLaradji/ssn 공식 구현 pytorch

Tasks

Binary ClassificationSecond-order methods

Similar Papers 제목 키워드 기반

Fast Black-box Variational Inference through Stochastic Trust-Region Optimization

2017-06-07 · NeurIPS 2017 12 · Jeffrey Regier, Michael. I. Jordan, Jon Mcauliffe

We introduce TrustVI, a fast second-order algorithm for black-box variational inference based on trust-region optimization and the reparameterization trick. At each iteration, TrustVI proposes and assesses a step based o…

Variational Inference

DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning

2025-10-02 · Hanyang Zhao, Dawen Liang, Wenpin Tang, David Yao 외 arxiv

We propose DiFFPO, Diffusion Fast and Furious Policy Optimization, a unified framework for training masked diffusion large language models (dLLMs) to reason not only better (furious), but also faster via reinforcement le…

Reinforcement Learning

Stochastic Newton and Cubic Newton Methods with Simple Local Linear-Quadratic Rates

2019-12-03 · Dmitry Kovalev, Konstantin Mishchenko, Peter Richtárik

We present two new remarkably simple stochastic second-order methods for minimizing the average of a very large number of sufficiently smooth and strongly convex functions. The first is a stochastic variant of Newton's m…

Second-order methods

Second-Order Stochastic Optimization for Machine Learning in Linear Time

2016-02-12 · Naman Agarwal, Brian Bullins, Elad Hazan

First-order stochastic methods are the state-of-the-art in large-scale machine learning optimization owing to efficient per-iteration complexity. Second-order methods, while able to provide faster convergence, have been …

BIG-bench Machine LearningSecond-order methodsStochastic Optimization

Farasa: A Fast and Furious Segmenter for Arabic

2016-06-01 · NAACL 2016 6 · Ahmed Abdelali, Kareem Darwish, Nadir Durrani, Hamdy Mubarak
Arabic Text DiacritizationInformation RetrievalMachine Translation