paper-with-me

홈 › Papers

The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton

2025-10-10 · Natalie Abreu, Nikhil Vyas, Sham Kakade, Depen Morwani arxiv

Recent efforts to accelerate LLM pretraining have focused on computationally-efficient approximations that exploit second-order structure. This raises a key question for large-scale training: how much performance is forfeited by these approximations? To probe this question, we establish a practical upper bound on iteration complexity by applying full Gauss-Newton (GN) preconditioning to transformer models of up to 150M parameters. Our experiments show that full GN updates yield substantial gains over existing optimizers, achieving a 5.4x reduction in training iterations compared to strong baselines like SOAP and Muon. Furthermore, we find that a precise layerwise GN preconditioner, which ignores cross-layer information, nearly matches the performance of the full GN method. Collectively, our results suggest: (1) the GN approximation is highly effective for preconditioning, implying higher-order loss terms may not be critical for convergence speed; (2) the layerwise Hessian structure contains sufficient information to achieve most of these potential gains; and (3) a significant performance gap exists between current approximate methods and an idealized layerwise oracle.

📄 PDF Abstract BibTeX arXiv:2510.09378

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SOUL: Unlocking the Power of Second-Order Optimization for LLM Unlearning

2024-04-28 · Jinghan Jia, Yihua Zhang, Yimeng Zhang, Jiancheng Liu 외

Large Language Models (LLMs) have highlighted the necessity of effective unlearning mechanisms to comply with data regulations and ethical AI practices. LLM unlearning aims at removing undesired data influences and assoc…

Stochastic Optimization

Avalon's Game of Thoughts: Battle Against Deception through Recursive Contemplation

2023-10-02 · Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi 외

Recent breakthroughs in large language models (LLMs) have brought remarkable success in the field of LLM-as-Agent. Nevertheless, a prevalent assumption is that the information processed by LLMs is consistently honest, ne…

Misinformation

Efficient Second-Order Neural Network Optimization via Adaptive Trust Region Methods

2024-10-03 · James Vo

Second-order optimization methods offer notable advantages in training deep neural networks by utilizing curvature information to achieve faster convergence. However, traditional second-order techniques are computational…

Computational Efficiency

Escaping Saddle Points in Nonconvex Minimax Optimization via Cubic-Regularized Gradient Descent-Ascent

2021-09-29 · Ziyi Chen, Qunwei Li, Yi Zhou

The gradient descent-ascent (GDA) algorithm has been widely applied to solve nonconvex minimax optimization problems. However, the existing GDA-type algorithms can only find first-order stationary points of the envelope …

VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models

2024-09-25 · Yifei Liu, Jicheng Wen, Yang Wang, Shengyu Ye 외

Scaling model size significantly challenges the deployment and inference of Large Language Models (LLMs). Due to the redundancy in LLM weights, recent research has focused on pushing weight-only quantization to extremely…

Quantization