paper-with-me

홈 › Papers

From SGD to Spectra: A Theory of Neural Network Weight Dynamics

2025-07-17 · Brian Richard Olsen, Sam Fatehmanesh, Frank Xiao, Adarsh Kumarappan, Anirudh Gajula arxiv

Deep neural networks have revolutionized machine learning, yet their training dynamics remain theoretically unclear-we develop a continuous-time, matrix-valued stochastic differential equation (SDE) framework that rigorously connects the microscopic dynamics of SGD to the macroscopic evolution of singular-value spectra in weight matrices. We derive exact SDEs showing that squared singular values follow Dyson Brownian motion with eigenvalue repulsion, and characterize stationary distributions as gamma-type densities with power-law tails, providing the first theoretical explanation for the empirically observed 'bulk+tail' spectral structure in trained networks. Through controlled experiments on transformer and MLP architectures, we validate our theoretical predictions and demonstrate quantitative agreement between SDE-based forecasts and observed spectral evolution, providing a rigorous foundation for understanding why deep learning works.

📄 PDF Abstract BibTeX arXiv:2507.12709

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer

2026-05-08 · Clarissa Lauditi, Cengiz Pehlevan, Blake Bordelon arxiv

We study the evolution of hidden-weight spectra in wide neural networks trained by (stochastic) gradient descent. We develop a two-level dynamical mean-field theory (DMFT) that jointly tracks bulk and outlier spectral dy…

An Analytical Theory of Power Law Spectral Bias in the Learning Dynamics of Diffusion Models

2025-03-05 · Binxu Wang

We developed an analytical framework for understanding how the learned distribution evolves during diffusion model training. Leveraging the Gaussian equivalence principle, we derived exact solutions for the gradient-flow…

Random matrix theory of sparse neuronal networks with heterogeneous timescales

2025-12-14 · Thiparat Chotibut, Oleg Evnin, Weerawit Horinouchi arxiv

Training recurrent neuronal networks consisting of excitatory (E) and inhibitory (I) units with additive noise for working memory computation slows and diversifies inhibitory timescales, leading to improved task performa…

Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws

2025-09-29 · Fabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani, Yizhou Xu 외 arxiv

Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the origin of these phenomena and their imp…

Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics

2026-07-12 · Byung Gyu Chae arxiv

Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood. Cognitive Field Theory predicts that learning reorganizes the collect…