paper-with-me

홈 › Papers

Enhance Curvature Information by Structured Stochastic Quasi-Newton Methods

2020-06-17 · CVPR 2021 1 · Ming-Han Yang, Dong Xu, Hongyu Chen, Zaiwen Wen, Mengyun Chen

In this paper, we consider stochastic second-order methods for minimizing a finite summation of nonconvex functions. One important key is to find an ingenious but cheap scheme to incorporate local curvature information. Since the true Hessian matrix is often a combination of a cheap part and an expensive part, we propose a structured stochastic quasi-Newton method by using partial Hessian information as much as possible. By further exploiting either the low-rank structure or the kronecker-product properties of the quasi-Newton approximations, the computation of the quasi-Newton direction is affordable. Global convergence to stationary point and local superlinear convergence rate are established under some mild assumptions. Numerical results on logistic regression, deep autoencoder networks and deep convolutional neural networks show that our proposed method is quite competitive to the state-of-the-art methods.

📄 PDF Abstract BibTeX arXiv:2006.09606

Code (0)

등록된 구현이 없습니다.

Tasks

Second-order methods

Methods 이 논문이 사용한 방법론

Solana Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

A Stochastic Quasi-Newton Method for Large-Scale Optimization

2014-01-27 · R. H. Byrd, S. L. Hansen, J. Nocedal, Y. Singer

The question of how to incorporate curvature information in stochastic approximation methods is challenging. The direct application of classical quasi- Newton updating techniques for deterministic optimization leads to n…

Symmetric Rank-One Quasi-Newton Methods for Deep Learning Using Cubic Regularization

2025-02-17 · Aditya Ranganath, Mukesh Singhal, Roummel Marcia

Stochastic gradient descent and other first-order variants, such as Adam and AdaGrad, are commonly used in the field of deep learning due to their computational efficiency and low-storage memory requirements. However, th…

Computational Efficiency

L-SR1 Adaptive Regularization by Cubics for Deep Learning

2021-09-29 · Aditya Ranganath, Mukesh Singhal, Roummel Marcia

Stochastic gradient descent and other first-order variants, such as Adam and AdaGrad, are commonly used in the field of deep learning due to their computational efficiency and low-storage memory requirements. However, th…

Computational EfficiencyDeep Learning

A Stochastic Variance Reduced Nesterov's Accelerated Quasi-Newton Method

2019-10-17 · Sota Yasuda, Shahrzad Mahboubi, S. Indrapriyadarsini, Hiroshi Ninomiya 외

Recently algorithms incorporating second order curvature information have become popular in training neural networks. The Nesterov's Accelerated Quasi-Newton (NAQ) method has shown to effectively accelerate the BFGS quas…

regression

adaQN: An Adaptive Quasi-Newton Algorithm for Training RNNs

2015-11-04 · Nitish Shirish Keskar, Albert S. Berahas

Recurrent Neural Networks (RNNs) are powerful models that achieve exceptional performance on several pattern recognition problems. However, the training of RNNs is a computationally difficult task owing to the well-known…

Language ModelingLanguage Modelling