paper-with-me

홈 › Papers

Research of Damped Newton Stochastic Gradient Descent Method for Neural Network Training

2021-03-31 · Jingcheng Zhou, Wei Wei, Zhiming Zheng

First-order methods like stochastic gradient descent(SGD) are recently the popular optimization method to train deep neural networks (DNNs), but second-order methods are scarcely used because of the overpriced computing cost in getting the high-order information. In this paper, we propose the Damped Newton Stochastic Gradient Descent(DN-SGD) method and Stochastic Gradient Descent Damped Newton(SGD-DN) method to train DNNs for regression problems with Mean Square Error(MSE) and classification problems with Cross-Entropy Loss(CEL), which is inspired by a proved fact that the hessian matrix of last layer of DNNs is always semi-definite. Different from other second-order methods to estimate the hessian matrix of all parameters, our methods just accurately compute a small part of the parameters, which greatly reduces the computational cost and makes convergence of the learning process much faster and more accurate than SGD. Several numerical experiments on real datesets are performed to verify the effectiveness of our methods for regression and classification problems.

📄 PDF Abstract BibTeX arXiv:2103.16764

Code (0)

등록된 구현이 없습니다.

Tasks

regressionSecond-order methods

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…

Similar Papers 제목 키워드 기반

Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent

2020-12-07 · Kangqiao Liu, Liu Ziyin, Masahito Ueda

In the vanishing learning rate regime, stochastic gradient descent (SGD) is now relatively well understood. In this work, we propose to study the basic properties of SGD and its variants in the non-vanishing learning rat…

Bayesian InferenceSecond-order methods

Error estimates between SGD with momentum and underdamped Langevin diffusion

2024-10-22 · Arnaud Guillin, Yu Wang, Lihu Xu, Haoran Yang

Stochastic gradient descent with momentum is a popular variant of stochastic gradient descent, which has recently been reported to have a close relationship with the underdamped Langevin diffusion. In this paper, we esta…

Efficient Numerical Algorithm for Large-Scale Damped Natural Gradient Descent

2023-10-26 · Yixiao Chen, Hao Xie, Han Wang

We propose a new algorithm for efficiently solving the damped Fisher matrix in large-scale scenarios where the number of parameters significantly exceeds the number of available samples. This problem is fundamental for n…

Modified Gauss-Newton Algorithms under Noise

2023-05-18 · Krishna Pillutla, Vincent Roulet, Sham Kakade, Zaid Harchaoui

Gauss-Newton methods and their stochastic version have been widely used in machine learning and signal processing. Their nonsmooth counterparts, modified Gauss-Newton or prox-linear algorithms, can lead to contrasting ou…

Structured Prediction

Randomness and Interpolation Improve Gradient Descent

2025-10-14 · Jiawen Li, Pascal Lefevre, Anwar Pp Abdul Majeed arxiv

Based on Stochastic Gradient Descent (SGD), the paper introduces two optimizers, named Interpolational Accelerating Gradient Descent (IAGD) as well as Noise-Regularized Stochastic Gradient Descent (NRSGD). IAGD leverages…