paper-with-me

홈 › Papers

Directional Convergence Near Small Initializations and Saddles in Two-Homogeneous Neural Networks

2024-02-14 · Akshay Kumar, Jarvis Haupt

This paper examines gradient flow dynamics of two-homogeneous neural networks for small initializations, where all weights are initialized near the origin. For both square and logistic losses, it is shown that for sufficiently small initializations, the gradient flow dynamics spend sufficient time in the neighborhood of the origin to allow the weights of the neural network to approximately converge in direction to the Karush-Kuhn-Tucker (KKT) points of a neural correlation function that quantifies the correlation between the output of the neural network and corresponding labels in the training data set. For square loss, it has been observed that neural networks undergo saddle-to-saddle dynamics when initialized close to the origin. Motivated by this, this paper also shows a similar directional convergence among weights of small magnitude in the neighborhood of certain saddle points.

📄 PDF Abstract BibTeX arXiv:2402.09226

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations

2024-03-12 · Akshay Kumar, Jarvis Haupt

This paper studies the gradient flow dynamics that arise when training deep homogeneous neural networks assumed to have locally Lipschitz gradients and an order of homogeneity strictly greater than two. It is shown here …

A Rigorous Framework for the Mean Field Limit of Multilayer Neural Networks

2020-01-30 · Phan-Minh Nguyen, Huy Tuan Pham

We develop a mathematically rigorous framework for multilayer neural networks in the mean field regime. As the network's widths increase, the network's learning trajectory is shown to be well captured by a meaningful and…

Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

2021-06-30 · Arthur Jacot, François Ged, Berfin Şimşek, Clément Hongler 외

The dynamics of Deep Linear Networks (DLNs) is dramatically affected by the variance $\sigma^2$ of the parameters at initialization $\theta_0$. For DLNs of width $w$, we show a phase transition w.r.t. the scaling $\gamma…

L2 Regularization

A proof of convergence for the gradient descent optimization method with random initializations in the training of neural networks with ReLU activation for piecewise linear target functions

2021-08-10 · Arnulf Jentzen, Adrian Riekert

Gradient descent (GD) type optimization methods are the standard instrument to train artificial neural networks (ANNs) with rectified linear unit (ReLU) activation. Despite the great success of GD type optimization metho…

Faster Directional Convergence of Linear Neural Networks under Spherically Symmetric Data

2021-12-01 · NeurIPS 2021 12 · Dachao Lin, Ruoyu Sun, Zhihua Zhang

In this paper, we study gradient methods for training deep linear neural networks with binary cross-entropy loss. In particular, we show global directional convergence guarantees from a polynomial rate to a linear rate f…