paper-with-me

홈 › Papers

Why Deep Jacobian Spectra Separate: Depth-Induced Scaling and Singular-Vector Alignment

2026-02-12 · Nathanaël Haas, François Gatine, Augustin M Cosse, Zied Bouraoui arxiv

Understanding why gradient-based training in deep networks exhibits strong implicit bias remains challenging, in part because tractable singular-value dynamics are typically available only for balanced deep linear models. We propose an alternative route based on two theoretically grounded and empirically testable signatures of deep Jacobians: depth-induced exponential scaling of ordered singular values and strong spectral separation. Adopting a fixed-gates view of piecewise-linear networks, where Jacobians reduce to products of masked linear maps within a single activation region, we prove the existence of Lyapunov exponents governing the top singular values at initialization, give closed-form expressions in a tractable masked model, and quantify finite-depth corrections. We further show that sufficiently strong separation forces singular-vector alignment in matrix products, yielding an approximately shared singular basis for intermediate Jacobians. Together, these results motivate an approximation regime in which singular-value dynamics become effectively decoupled, mirroring classical balanced deep-linear analyses without requiring balancing. Experiments in fixed-gates settings validate the predicted scaling, alignment, and resulting dynamics, supporting a mechanistic account of emergent low-rank Jacobian structure as a driver of implicit bias.

📄 PDF Abstract BibTeX arXiv:2602.12384

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On the Stability of the Jacobian Matrix in Deep Neural Networks

2025-06-10 · Benjamin Dadoun, Soufiane Hayou, Hanan Salam, Mohamed El Amine Seddik 외

Deep neural networks are known to suffer from exploding or vanishing gradients as depth increases, a phenomenon closely tied to the spectral behavior of the input-output Jacobian. Prior work has identified critical initi…

Renormalizable Spectral-Shell Dynamics as the Origin of Neural Scaling Laws

2025-12-11 · Yizhou Zhang arxiv

Neural scaling laws and double-descent phenomena suggest that deep-network training obeys a simple macroscopic structure despite highly nonlinear optimization dynamics. We derive such structure directly from gradient des…

The Emergence of Spectral Universality in Deep Networks

2018-02-27 · Jeffrey Pennington, Samuel S. Schoenholz, Surya Ganguli

Recent work has shown that tight concentration of the entire spectrum of singular values of a deep network's input-output Jacobian around one at initialization can speed up learning by orders of magnitude. Therefore, to …

Stabilizing Recurrent Dynamics for Test-Time Scalable Latent Reasoning in Looped Language Models

2026-05-26 · Xiao-Wen Yang, Ziyu Han, Xi-Hua Zhang, Wen-Da Wei 외 arxiv

Looped Language Models (LoopLMs) enable efficient latent reasoning through depth recurrence, yet exhibit unreliable test-time scaling behavior: performance often peaks at a certain iteration depth and then collapses with…

Mathematical Reasoning

Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics

2026-07-12 · Byung Gyu Chae arxiv

Large language models exhibit remarkable emergent behaviors, yet the physical mechanism governing their collective dynamics remains poorly understood. Cognitive Field Theory predicts that learning reorganizes the collect…