paper-with-me

Papers

Large-width asymptotics for ReLU neural networks with $α$-Stable initializations

2022-06-16 · Stefano Favaro, Sandra Fortini, Stefano Peluchetti

There is a recent and growing literature on large-width asymptotic properties of Gaussian neural networks (NNs), namely NNs whose weights are initialized as Gaussian distributions. Two popular problems are: i) the study of the large-width distributions of NNs, which characterizes the infinitely wide limit of a rescaled NN in terms of a Gaussian stochastic process; ii) the study of the large-width training dynamics of NNs, which characterizes the infinitely wide dynamics in terms of a deterministic kernel, referred to as the neural tangent kernel (NTK), and shows that, for a sufficiently large width, the gradient descent achieves zero training error at a linear rate. In this paper, we consider these problems for $\alpha$-Stable NNs, namely NNs whose weights are initialized as $\alpha$-Stable distributions with $\alpha\in(0,2]$. First, for $\alpha$-Stable NNs with a ReLU activation function, we show that if the NN's width goes to infinity then a rescaled NN converges weakly to an $\alpha$-Stable stochastic process. As a difference with respect to the Gaussian setting, our result shows that the choice of the activation function affects the scaling of the NN, that is: to achieve the infinitely wide $\alpha$-Stable process, the ReLU activation requires an additional logarithmic term in the scaling with respect to sub-linear activations. Then, we study the large-width training dynamics of $\alpha$-Stable ReLU-NNs, characterizing the infinitely wide dynamics in terms of a random kernel, referred to as the $\alpha$-Stable NTK, and showing that, for a sufficiently large width, the gradient descent achieves zero training error at a linear rate. The randomness of the $\alpha$-Stable NTK is a further difference with respect to the Gaussian setting, that is: within the $\alpha$-Stable setting, the randomness of the NN at initialization does not vanish in the large-width regime of the training.

📄 PDF Abstract BibTeX arXiv:2206.08065

Code (0)

등록된 구현이 없습니다.

Tasks

regression

Methods 이 논문이 사용한 방법론

NTK 설명 없음

Similar Papers 제목 키워드 기반

On the asymptotics of wide networks with polynomial activations

2020-06-11 · Kyle Aitken, Guy Gur-Ari

We consider an existing conjecture addressing the asymptotic behavior of neural networks in the large width limit. The results that follow from this conjecture include tight bounds on the behavior of wide networks during…

A proof of convergence for the gradient descent optimization method with random initializations in the training of neural networks with ReLU activation for piecewise linear target functions

2021-08-10 · Arnulf Jentzen, Adrian Riekert

Gradient descent (GD) type optimization methods are the standard instrument to train artificial neural networks (ANNs) with rectified linear unit (ReLU) activation. Despite the great success of GD type optimization metho…

Optimized Weight Initialization on the Stiefel Manifold for Deep ReLU Neural Networks

2025-08-30 · Hyungu Lee, Taehyeong Kim, Hayoung Choi arxiv

Stable and efficient training of ReLU networks with large depth is highly sensitive to weight initialization. Improper initialization can cause permanent neuron inactivation dying ReLU and exacerbate gradient instability…

Gradient Dynamics of Shallow Univariate ReLU Networks

2019-06-18 · NeurIPS 2019 12 · Francis Williams, Matthew Trager, Claudio Silva, Daniele Panozzo 외

We present a theoretical and empirical study of the gradient dynamics of overparameterized shallow ReLU networks with one-dimensional input, solving least-squares interpolation. We show that the gradient dynamics of such…

Deep Stable neural networks: large-width asymptotics and convergence rates

2021-08-02 · Stefano Favaro, Sandra Fortini, Stefano Peluchetti

In modern deep learning, there is a recent and growing literature on the interplay between large-width asymptotic properties of deep Gaussian neural networks (NNs), i.e. deep NNs with Gaussian-distributed weights, and Ga…

Bayesian Inference