paper-with-me

홈 › Papers

Deep Stable neural networks: large-width asymptotics and convergence rates

2021-08-02 · Stefano Favaro, Sandra Fortini, Stefano Peluchetti

In modern deep learning, there is a recent and growing literature on the interplay between large-width asymptotic properties of deep Gaussian neural networks (NNs), i.e. deep NNs with Gaussian-distributed weights, and Gaussian stochastic processes (SPs). Such an interplay has proved to be critical in Bayesian inference under Gaussian SP priors, kernel regression for infinitely wide deep NNs trained via gradient descent, and information propagation within infinitely wide NNs. Motivated by empirical analyses that show the potential of replacing Gaussian distributions with Stable distributions for the NN's weights, in this paper we present a rigorous analysis of the large-width asymptotic behaviour of (fully connected) feed-forward deep Stable NNs, i.e. deep NNs with Stable-distributed weights. We show that as the width goes to infinity jointly over the NN's layers, i.e. the `joint growth" setting, a rescaled deep Stable NN converges weakly to a Stable SP whose distribution is characterized recursively through the NN's layers. Because of the non-triangular structure of the NN, this is a non-standard asymptotic problem, to which we propose an inductive approach of independent interest. Then, we establish sup-norm convergence rates of the rescaled deep Stable NN to the Stable SP, under the joint growth" and a sequential growth" of the width over the NN's layers. Such a result provides the difference between the joint growth" and the sequential growth" settings, showing that the former leads to a slower rate than the latter, depending on the depth of the layer and the number of inputs of the NN. Our work extends some recent results on infinitely wide limits for deep Gaussian NNs to the more general deep Stable NNs, providing the first result on convergence rates in the `joint growth" setting.

📄 PDF Abstract BibTeX arXiv:2108.02316

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian Inference

Similar Papers 제목 키워드 기반

Large-width asymptotics for ReLU neural networks with $α$-Stable initializations

2022-06-16 · Stefano Favaro, Sandra Fortini, Stefano Peluchetti

There is a recent and growing literature on large-width asymptotic properties of Gaussian neural networks (NNs), namely NNs whose weights are initialized as Gaussian distributions. Two popular problems are: i) the study …

regression

Higher-order Refinements of Small Bandwidth Asymptotics for Density-Weighted Average Derivative Estimators

2022-12-31 · Matias D. Cattaneo, Max H. Farrell, Michael Jansson, Ricardo Masini

The density weighted average derivative (DWAD) of a regression function is a canonical parameter of interest in economics. Classical first-order large sample distribution theory for kernel-based DWAD estimators relies on…

Phase Diagram of Dropout for Two-Layer Neural Networks in the Mean-Field Regime

2025-10-08 · Lénaïc Chizat, Pierre Marion, Yerkin Yesbay arxiv

Dropout is a standard training technique for neural networks that consists of randomly deactivating units at each step of their gradient-based training. It is known to improve performance in many settings, including in t…

Large-width functional asymptotics for deep Gaussian neural networks

2021-02-20 · ICLR 2021 1 · Daniele Bracale, Stefano Favaro, Sandra Fortini, Stefano Peluchetti

In this paper, we consider fully connected feed-forward deep neural networks where weights and biases are independent and identically distributed according to Gaussian distributions. Extending previous results (Matthews …

Gaussian Processes

Convergence rates for Poisson learning to a Poisson equation with measure data

2024-07-09 · Leon Bungert, Jeff Calder, Max Mihailescu, Kodjo Houssou 외

In this paper we prove discrete to continuum convergence rates for Poisson Learning, a graph-based semi-supervised learning algorithm that is based on solving the graph Poisson equation with a source term consisting of a…