paper-with-me

홈 › Papers

Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations

2024-03-12 · Akshay Kumar, Jarvis Haupt

This paper studies the gradient flow dynamics that arise when training deep homogeneous neural networks assumed to have locally Lipschitz gradients and an order of homogeneity strictly greater than two. It is shown here that for sufficiently small initializations, during the early stages of training, the weights of the neural network remain small in (Euclidean) norm and approximately converge in direction to the Karush-Kuhn-Tucker (KKT) points of the recently introduced neural correlation function. Additionally, this paper also studies the KKT points of the neural correlation function for feed-forward networks with (Leaky) ReLU and polynomial (Leaky) ReLU activations, deriving necessary and sufficient conditions for rank-one KKT points.

📄 PDF Abstract BibTeX arXiv:2403.08121

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Directional Convergence Near Small Initializations and Saddles in Two-Homogeneous Neural Networks

2024-02-14 · Akshay Kumar, Jarvis Haupt

This paper examines gradient flow dynamics of two-homogeneous neural networks for small initializations, where all weights are initialized near the origin. For both square and logistic losses, it is shown that for suffic…

Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks

2025-02-22 · Yuhang Cai, Kangjie Zhou, Jingfeng Wu, Song Mei 외

We establish the asymptotic implicit bias of gradient descent (GD) for generic non-homogeneous deep networks under exponential loss. Specifically, we characterize three key properties of GD iterates starting from a suffi…

A Rigorous Framework for the Mean Field Limit of Multilayer Neural Networks

2020-01-30 · Phan-Minh Nguyen, Huy Tuan Pham

We develop a mathematically rigorous framework for multilayer neural networks in the mean field regime. As the network's widths increase, the network's learning trajectory is shown to be well captured by a meaningful and…

Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks

2020-01-16 · ICLR 2020 1 · Wei Hu, Lechao Xiao, Jeffrey Pennington

The selection of initial parameter values for gradient-based optimization of deep neural networks is one of the most impactful hyperparameter choices in deep learning systems, affecting both convergence times and model p…

A Note on the Global Convergence of Multilayer Neural Networks in the Mean Field Regime

2020-06-16 · Huy Tuan Pham, Phan-Minh Nguyen

In a recent work, we introduced a rigorous framework to describe the mean field limit of the gradient-based learning dynamics of multilayer neural networks, based on the idea of a neuronal embedding. There we also proved…

Diversity