paper-with-me

홈 › Papers

Exact Gauss-Newton Optimization for Training Deep Neural Networks

2024-05-23 · Mikalai Korbit, Adeyemi D. Adeoye, Alberto Bemporad, Mario Zanon

We present EGN, a stochastic second-order optimization algorithm that combines the generalized Gauss-Newton (GN) Hessian approximation with low-rank linear algebra to compute the descent direction. Leveraging the Duncan-Guttman matrix identity, the parameter update is obtained by factorizing a matrix which has the size of the mini-batch. This is particularly advantageous for large-scale machine learning problems where the dimension of the neural network parameter vector is several orders of magnitude larger than the batch size. Additionally, we show how improvements such as line search, adaptive regularization, and momentum can be seamlessly added to EGN to further accelerate the algorithm. Moreover, under mild assumptions, we prove that our algorithm converges to an $\epsilon$-stationary point at a linear rate. Finally, our numerical experiments demonstrate that EGN consistently exceeds, or at most matches the generalization performance of well-tuned SGD, Adam, and SGN optimizers across various supervised and reinforcement learning tasks.

📄 PDF Abstract BibTeX arXiv:2405.14402

Code (1)

cor3bit/somax 공식 구현 jax

Methods 이 논문이 사용한 방법론

SGD Stochastic Gradient Descent is an iterative optimization technique that uses minibatches of data to form an expectation of the gradient, rather than the full gradient using…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Newton-LESS: Sparsification without Trade-offs for the Sketched Newton Update

2021-07-15 · NeurIPS 2021 12 · Michał Dereziński, Jonathan Lacotte, Mert Pilanci, Michael W. Mahoney

In second-order optimization, a potential bottleneck can be computing the Hessian matrix of the optimized function at every iteration. Randomized sketching has emerged as a powerful technique for constructing estimates o…

Frugality in second-order optimization: floating-point approximations for Newton's method

2025-11-20 · Giuseppe Carrino, Elena Loli Piccolomini, Elisa Riccietti, Theo Mary arxiv

Minimizing loss functions is central to machine-learning training. Although first-order methods dominate practical applications, higher-order techniques such as Newton's method can deliver greater accuracy and faster con…

Dual Natural Gradient Descent for Scalable Training of Physics-Informed Neural Networks

2025-05-27 · Anas Jnini, Flavio Vella

Natural-gradient methods markedly accelerate the training of Physics-Informed Neural Networks (PINNs), yet their Gauss--Newton update must be solved in the parameter space, incurring a prohibitive $O(n^3)$ time complexit…

GPU

Newton-MR: Inexact Newton Method With Minimum Residual Sub-problem Solver

2018-09-30 · Fred Roosta, Yang Liu, Peng Xu, Michael W. Mahoney

We consider a variant of inexact Newton Method, called Newton-MR, in which the least-squares sub-problems are solved approximately using Minimum Residual method. By construction, Newton-MR can be readily applied for unco…

Gauss-Newton Natural Gradient Descent for Shape Learning

2026-01-24 · James King, Arturs Berzins, Siddhartha Mishra, Marius Zeinhofer arxiv

We explore the use of the Gauss-Newton method for optimization in shape learning, including implicit neural surfaces and geometry-informed neural networks. The method addresses key challenges in shape learning, such as t…