paper-with-me

홈 › Papers

Spectral Asymptotics of Neural Network Loss Landscapes: An Exact Decomposition of the Curvature Exponent

2026-05-22 · Anherutowa Calvo arxiv

The curvature exponent $α$ in $h_k \propto σ_k^α$ -- governing how Hessian eigenvalues scale with gradient singular values -- varies systematically across layer types ($α\approx 2$ for convolutions, $\approx 1$ for transformer attention, $< 1$ for MLP up-projections). Why? We prove the Spectral Alignment Decomposition: $α= 2 + d\logΦ_k / d\logσ_k$, where $Φ_k$ measures alignment between Kronecker factor eigenbases and gradient singular directions. This reduces "why does $α$ vary?" to a geometric question we answer for LayerNorm, residual connections, and softmax heads. The decomposition implies a spectral transfer identity $s = αγ$ linking curvature exponent, effective gradient rank-decay $γ$, and Hessian decay exponent $s$. The identity is algebraic; its empirical content is that $α$ and $γ$, fit on independent data (HVPs vs. SVD), recover $s$ to ~2% median error across 93 layers, five architectures, and three datasets -- with no free parameters. A zeta-function bound on participation ratio shows curvature concentrates onto effectively one direction per layer. As a proof of concept, we derive the architecture-adaptive preconditioner $T(σ;α)$ and show that Spectral Newton -- implementing $T$ in the gradient singular basis -- outperforms AdamW on vision benchmarks where $α\approx 2$.

📄 PDF Abstract BibTeX arXiv:2606.02596

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Precise Asymptotics for Spectral Methods in Mixed Generalized Linear Models

2022-11-21 · Yihan Zhang, Marco Mondelli, Ramji Venkataramanan

In a mixed generalized linear model, the objective is to learn multiple signals from unlabeled observations: each sample comes from exactly one signal, but it is not known which one. We consider the prototypical problem …

Retrieval

Kähler landscapes for complex neural network descents and guarantees including a search and destroy of the Calabi-Yau manifold

2026-08-20 · Andrew Gracyk arxiv

We study landscapes for complex-parameterized networks. Our approach is motivated with an information-theoretic manifold perspective of the parameter and via classical optimization guarantees although of complex geometri…

Evaluating Loss Landscapes from a Topology Perspective

2024-11-14 · Tiankai Xie, Caleb Geniesse, Jiaqing Chen, Yaoqing Yang 외

Characterizing the loss of a neural network with respect to model parameters, i.e., the loss landscape, can provide valuable insights into properties of that model. Various methods for visualizing loss landscapes have be…

Topological Data Analysis

On magnitude, asymptotics and duration of drawdowns for L\'{e}vy models

2016-09-30

This paper considers magnitude, asymptotics and duration of drawdowns for some L\'{e}vy processes. First, we revisit some existing results on the magnitude of drawdowns for spectrally negative L\'{e}vy processes using an…

Management

A simpler spectral approach for clustering in directed networks

2021-02-05 · Simon Coste, Ludovic Stephan

We study the task of clustering in directed networks. We show that using the eigenvalue/eigenvector decomposition of the adjacency matrix is simpler than all common methods which are based on a combination of data regula…

Clustering