paper-with-me

Papers

Spectral-factorized Positive-definite Curvature Learning for NN Training

2025-02-10 · Wu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae, Richard E. Turner, Roger B. Grosse

Many training methods, such as Adam(W) and Shampoo, learn a positive-definite curvature matrix and apply an inverse root before preconditioning. Recently, non-diagonal training methods, such as Shampoo, have gained significant attention; however, they remain computationally inefficient and are limited to specific types of curvature information due to the costly matrix root computation via matrix decomposition. To address this, we propose a Riemannian optimization approach that dynamically adapts spectral-factorized positive-definite curvature estimates, enabling the efficient application of arbitrary matrix roots and generic curvature learning. We demonstrate the efficacy and versatility of our approach in positive-definite matrix optimization and covariance adaptation for gradient-free optimization, as well as its efficiency in curvature learning for neural net training.

📄 PDF Abstract BibTeX arXiv:2502.06268

Code (0)

등록된 구현이 없습니다.

Tasks

Riemannian optimization

Similar Papers 제목 키워드 기반

On the Convex Behavior of Deep Neural Networks in Relation to the Layers' Width

2020-01-14 · ICML Workshop Deep_Phenomen 2019 6 · Etai Littwin, Lior Wolf

The Hessian of neural networks can be decomposed into a sum of two matrices: (i) the positive semidefinite generalized Gauss-Newton matrix G, and (ii) the matrix H containing negative eigenvalues. We observe that for wid…

Relation

Geodesic Exponential Kernels: When Curvature and Linearity Conflict

2014-11-02 · CVPR 2015 6 · Aasa Feragen, Francois Lauze, Søren Hauberg

We consider kernel methods on general geodesic metric spaces and provide both negative and positive results. First we show that the common Gaussian kernel can only be generalized to a positive definite kernel on a geodes…

Fast Gauss-Newton for Multiclass Cross-Entropy

2026-05-07 · Mikalai Korbit, Mario Zanon arxiv

In multiclass softmax cross-entropy, the full generalized Gauss-Newton (GGN) curvature couples all output logits through the softmax covariance, making curvature-vector products harder to scale as the number of classes g…

Binary Classification

DynMuon: A Dynamic Spectral Shaping View of Muon

2026-05-16 · Fangzhou Wu, Rikhav Shah, Sandeep Silwal, Qiuyi Zhang arxiv

In recent years, Muon has emerged as the dominant method for training large language models, and transformers more broadly. The essential difference, when compared to standard gradient descent methods, is to replace the …

From Non-Convex Self-Concordant Regularization to Scalable Quasi-Newton Training of PINNs

2026-08-04 · Chenhao Si, Kang An, Shiqian Ma, Ming Yan arxiv

Physics-informed neural networks (PINNs) often require high-accuracy quasi-Newton refinement to obtain reliable partial differential equation solutions, but their residual objectives can exhibit indefinite, nearly singul…