paper-with-me

홈 › Papers

$α$-Stable convergence of heavy-tailed infinitely-wide neural networks

2021-06-18 · Paul Jung, Hoil Lee, Jiho Lee, Hongseok Yang

We consider infinitely-wide multi-layer perceptrons (MLPs) which are limits of standard deep feed-forward neural networks. We assume that, for each layer, the weights of an MLP are initialized with i.i.d. samples from either a light-tailed (finite variance) or heavy-tailed distribution in the domain of attraction of a symmetric $\alpha$-stable distribution, where $\alpha\in(0,2]$ may depend on the layer. For the bias terms of the layer, we assume i.i.d. initializations with a symmetric $\alpha$-stable distribution having the same $\alpha$ parameter of that layer. We then extend a recent result of Favaro, Fortini, and Peluchetti (2020), to show that the vector of pre-activation values at all nodes of a given hidden layer converges in the limit, under a suitable scaling, to a vector of i.i.d. random variables with symmetric $\alpha$-stable distributions.

📄 PDF Abstract BibTeX arXiv:2106.11064

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Infinitely wide limits for deep Stable neural networks: sub-linear, linear and super-linear activation functions

2023-04-08 · Alberto Bordino, Stefano Favaro, Sandra Fortini

There is a growing literature on the study of large-width properties of deep Gaussian neural networks (NNs), i.e. deep NNs with Gaussian-distributed parameters or weights, and Gaussian stochastic processes. Motivated by …

Deep Stable neural networks: large-width asymptotics and convergence rates

2021-08-02 · Stefano Favaro, Sandra Fortini, Stefano Peluchetti

In modern deep learning, there is a recent and growing literature on the interplay between large-width asymptotic properties of deep Gaussian neural networks (NNs), i.e. deep NNs with Gaussian-distributed weights, and Ga…

Bayesian Inference

On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks

2019-11-29 · Umut Şimşekli, Mert Gürbüzbalaban, Thanh Huy Nguyen, Gaël Richard 외

The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the \emph{classical} central limit theorem (CLT) kicks in. This assu…

Scale Mixtures of Neural Network Gaussian Processes

2021-07-03 · ICLR 2022 4 · Hyungi Lee, Eunggu Yun, Hongseok Yang, Juho Lee

Recent works have revealed that infinitely-wide feed-forward or recurrent neural networks of any architecture correspond to Gaussian processes referred to as Neural Network Gaussian Processes (NNGPs). While these works h…

Gaussian Processes

Convergence of Heavy-Tailed Hawkes Processes and the Microstructure of Rough Volatility

2023-12-14 · Ulrich Horst, Wei Xu, Rouyi Zhang

We establish the weak convergence of the intensity of a nearly-unstable Hawkes process with heavy-tailed kernel. Our result is used to derive a scaling limit for a financial market model where orders to buy or sell an as…