paper-with-me

홈 › Papers

On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature

2026-02-05 · Yikuan Zhang, Ning Yang, Yuhai Tu arxiv

Stochastic Gradient Descent (SGD) introduces anisotropic noise that is correlated with the local curvature of the loss landscape, thereby biasing optimization toward flat minima. Prior work often assumes an equivalence between the Fisher Information Matrix and the Hessian for negative log-likelihood losses, leading to the claim that the SGD noise covariance $\mathbf{C}$ is proportional to the Hessian $\mathbf{H}$. We show that this assumption holds only under restrictive conditions that are typically violated in deep neural networks. Using the recently discovered Activity--Weight Duality, we find a more general relationship agnostic to the specific loss formulation, showing that $\mathbf{C} \propto \mathbb{E}_p[\mathbf{h}_p^2]$, where $\mathbf{h}_p$ denotes the per-sample Hessian with $\mathbf{H} = \mathbb{E}_p[\mathbf{h}_p]$. As a consequence, $\mathbf{C}$ and $\mathbf{H}$ commute approximately rather than coincide exactly. We further find that, within the analyzed fully connected layers, their diagonal elements follow per-layer empirical power laws $C_{ii} \propto H_{ii}^γ$, with layer-dependent fitted exponents bounded by $1 \leq γ\leq 2$. Experiments across datasets, architectures, and loss functions support the resulting layerwise bounds, providing a unified characterization of the noise-curvature relationship in deep learning.

📄 PDF Abstract BibTeX arXiv:2602.05600

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Nighttime Light, Superlinear Growth, and Economic Inequalities at the Country Level

2018-10-30

Research has highlighted relationships between size and scaled growth across a large variety of biological and social organisms, ranging from bacteria, through animals and plants, to cities an companies. Yet, heretofore,…

Hessian Averaging in Stochastic Newton Methods Achieves Superlinear Convergence

2022-04-20 · Sen Na, Michał Dereziński, Michael W. Mahoney

We consider minimizing a smooth and strongly convex objective function using a stochastic Newton method. At each iteration, the algorithm is given an oracle access to a stochastic estimate of the Hessian matrix. The orac…

Quasi-potential as an implicit regularizer for the loss function in the stochastic gradient descent

2019-01-18 · Wenqing Hu, Zhanxing Zhu, Haoyi Xiong, Jun Huan

We interpret the variational inference of the Stochastic Gradient Descent (SGD) as minimizing a new potential function named the \textit{quasi-potential}. We analytically construct the quasi-potential function in the cas…

RelationVariational Inference

Stochastic Newton Proximal Extragradient Method

2024-06-03 · Ruichen Jiang, Michał Dereziński, Aryan Mokhtari

Stochastic second-order methods achieve fast local convergence in strongly convex optimization by using noisy Hessian estimates to precondition the gradient. However, these methods typically reach superlinear convergence…

Second-order methods

Shift-Curvature, SGD, and Generalization

2021-08-21 · Arwen V. Bradley, Carlos Alberto Gomez-Uribe, Manish Reddy Vuyyuru

A longstanding debate surrounds the related hypotheses that low-curvature minima generalize better, and that SGD discourages curvature. We offer a more complete and nuanced view in support of both. First, we show that cu…