paper-with-me

홈 › Papers

Exact Stochastic Second Order Deep Learning

2021-04-08 · Fares B. Mehouachi, Chaouki Kasmi

Optimization in Deep Learning is mainly dominated by first-order methods which are built around the central concept of backpropagation. Second-order optimization methods, which take into account the second-order derivatives are far less used despite superior theoretical properties. This inadequacy of second-order methods stems from its exorbitant computational cost, poor performance, and the ineluctable non-convex nature of Deep Learning. Several attempts were made to resolve the inadequacy of second-order optimization without reaching a cost-effective solution, much less an exact solution. In this work, we show that this long-standing problem in Deep Learning could be solved in the stochastic case, given a suitable regularization of the neural network. Interestingly, we provide an expression of the stochastic Hessian and its exact eigenvalues. We provide a closed-form formula for the exact stochastic second-order Newton direction, we solve the non-convexity issue and adjust our exact solution to favor flat minima through regularization and spectral adjustment. We test our exact stochastic second-order method on popular datasets and reveal its adequacy for Deep Learning.

📄 PDF Abstract BibTeX arXiv:2104.03804

Code (0)

등록된 구현이 없습니다.

Tasks

Deep LearningSecond-order methods

Similar Papers 제목 키워드 기반

Stochastic Second-order Methods for Non-convex Optimization with Inexact Hessian and Gradient

2018-09-26 · Liu Liu, Xuanqing Liu, Cho-Jui Hsieh, DaCheng Tao

Trust region and cubic regularization methods have demonstrated good performance in small scale non-convex optimization, showing the ability to escape from saddle points. Each iteration of these methods involves computat…

Second-order methods

Exact Stochastic Newton Method for Deep Learning: the feedforward networks case.

2021-09-29 · Fares B. Mehouachi, Chaouki Kasmi

The inclusion of second-order information into Deep Learning optimization has drawn consistent interest as a way forward to improve upon gradient descent methods. Estimating the second-order update is often convoluted an…

Deep Learning

Asymptotic Implied Volatility at the Second Order with Application to the SABR Model

2016-05-17

We provide a general method to compute a Taylor expansion in time of implied volatility for stochastic volatility models, using a heat kernel expansion. Beyond the order 0 implied volatility which is already known, we co…

Stochastic Optimization for Non-convex Problem with Inexact Hessian Matrix, Gradient, and Function

2023-10-18 · Liu Liu, Xuanqing Liu, Cho-Jui Hsieh, DaCheng Tao

Trust-region (TR) and adaptive regularization using cubics (ARC) have proven to have some very appealing theoretical properties for non-convex optimization by concurrently computing function value, gradient, and Hessian …

ARCSecond-order methodsStochastic Optimization

A Novel Fast Exact Subproblem Solver for Stochastic Quasi-Newton Cubic Regularized Optimization

2022-04-19 · Jarad Forristal, Joshua Griffin, Wenwen Zhou, Seyedalireza Yektamaram

In this work we describe an Adaptive Regularization using Cubics (ARC) method for large-scale nonconvex unconstrained optimization using Limited-memory Quasi-Newton (LQN) matrices. ARC methods are a relatively new family…

ARCSecond-order methods