paper-with-me

Papers

Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs

2022-06-02 · Etienne Boursier, Loucas Pillaud-Vivien, Nicolas Flammarion

The training of neural networks by gradient descent methods is a cornerstone of the deep learning revolution. Yet, despite some recent progress, a complete theory explaining its success is still missing. This article presents, for orthogonal input vectors, a precise description of the gradient flow dynamics of training one-hidden layer ReLU neural networks for the mean squared error at small initialisation. In this setting, despite non-convexity, we show that the gradient flow converges to zero loss and characterise its implicit bias towards minimum variation norm. Furthermore, some interesting phenomena are highlighted: a quantitative description of the initial alignment phenomenon and a proof that the process follows a specific saddle to saddle dynamics.

📄 PDF Abstract BibTeX arXiv:2206.00939

Code (1)

eboursier/gfdynamics 공식 구현

Similar Papers 제목 키워드 기반

Gradient Dynamics of Shallow Univariate ReLU Networks

2019-06-18 · NeurIPS 2019 12 · Francis Williams, Matthew Trager, Claudio Silva, Daniele Panozzo 외

We present a theoretical and empirical study of the gradient dynamics of overparameterized shallow ReLU networks with one-dimensional input, solving least-squares interpolation. We show that the gradient dynamics of such…

Magnitude and Angle Dynamics in Training Single ReLU Neurons

2022-09-27 · Sangmin Lee, Byeongsu Sim, Jong Chul Ye

To understand learning the dynamics of deep ReLU networks, we investigate the dynamic system of gradient flow $w(t)$ by decomposing it to magnitude $w(t)$ and angle $\phi(t):= \pi - \theta(t) $ components. In particular,…

Support Vectors and Gradient Dynamics of Single-Neuron ReLU Networks

2022-02-11 · Sangmin Lee, Byeongsu Sim, Jong Chul Ye

Understanding implicit bias of gradient descent for generalization capability of ReLU networks has been an important research topic in machine learning research. Unfortunately, even for a single ReLU neuron trained with …

Gradient descent provably escapes saddle points in the training of shallow ReLU networks

2022-08-03 · Patrick Cheridito, Arnulf Jentzen, Florian Rossmannek

Dynamical systems theory has recently been applied in optimization to prove that gradient descent algorithms bypass so-called strict saddle points of the loss function. However, in many modern machine learning applicatio…

Quantitative Convergence of Wasserstein Gradient Flows of Kernel Mean Discrepancies

2026-03-02 · Lénaïc Chizat, Maria Colombo, Roberto Colombo, Xavier Fernández-Real arxiv

We study the quantitative convergence of Wasserstein gradient flows of Kernel Mean Discrepancy (KMD) (also known as Maximum Mean Discrepancy (MMD)) functionals. Our setting covers in particular the training dynamics of s…