paper-with-me

홈 › Papers

Towards Understanding Gradient Flow Dynamics of Homogeneous Neural Networks Beyond the Origin

2025-02-21 · Akshay Kumar, Jarvis Haupt

Recent works exploring the training dynamics of homogeneous neural network weights under gradient flow with small initialization have established that in the early stages of training, the weights remain small and near the origin, but converge in direction. Building on this, the current paper studies the gradient flow dynamics of homogeneous neural networks with locally Lipschitz gradients, after they escape the origin. Insights gained from this analysis are used to characterize the first saddle point encountered by gradient flow after escaping the origin. Also, it is shown that for homogeneous feed-forward neural networks, under certain conditions, the sparsity structure emerging among the weights before the escape is preserved after escaping the origin and until reaching the next saddle point.

📄 PDF Abstract BibTeX arXiv:2502.15952

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Asymmetric Maximum Margin Bias of Quasi-Homogeneous Neural Networks

2022-10-07 · Daniel Kunin, Atsushi Yamamura, Chao Ma, Surya Ganguli

In this work, we explore the maximum-margin bias of quasi-homogeneous neural networks trained with gradient flow on an exponential loss and past a point of separability. We introduce the class of quasi-homogeneous models…

Directional Convergence Near Small Initializations and Saddles in Two-Homogeneous Neural Networks

2024-02-14 · Akshay Kumar, Jarvis Haupt

This paper examines gradient flow dynamics of two-homogeneous neural networks for small initializations, where all weights are initialized near the origin. For both square and logistic losses, it is shown that for suffic…

Fisher information dissipation for time inhomogeneous stochastic differential equations

2024-02-01 · Qi Feng, Xinzhe Zuo, Wuchen Li

We provide a Lyapunov convergence analysis for time-inhomogeneous variable coefficient stochastic differential equations (SDEs). Three typical examples include overdamped, irreversible drift, and underdamped Langevin dyn…

Implicit Bias of Mirror Flow in Homogeneous Neural Networks: Sparse and Dense Feature Learning

2026-05-19 · Tom Jacobs, Guido Montufar arxiv

We study the max-margin solutions reached by mirror flow in deep neural networks with homogeneous activation functions. Extending classical results on gradient flow, we derive a novel balance equation for mirror flow fro…

The late-stage training dynamics of (stochastic) subgradient descent on homogeneous neural networks

2025-02-08 · Sholom Schechtman, Nicolas Schreuder

We analyze the implicit bias of constant step stochastic subgradient descent (SGD). We consider the setting of binary classification with homogeneous neural networks - a large class of deep neural networks with ReLU-type…

Binary Classification