paper-with-me

Papers

Infinitely Deep Infinite-Width Networks

2019-05-01 · ICLR 2019 5 · Jovana Mitrovic, Peter Wirnsberger, Charles Blundell, Dino Sejdinovic, Yee Whye Teh

Infinite-width neural networks have been extensively used to study the theoretical properties underlying the extraordinary empirical success of standard, finite-width neural networks. Nevertheless, until now, infinite-width networks have been limited to at most two hidden layers. To address this shortcoming, we study the initialisation requirements of these networks and show that the main challenge for constructing them is defining the appropriate sampling distributions for the weights. Based on these observations, we propose a principled approach to weight initialisation that correctly accounts for the functional nature of the hidden layer activations and facilitates the construction of arbitrarily many infinite-width layers, thus enabling the construction of arbitrarily deep infinite-width networks. The main idea of our approach is to iteratively reparametrise the hidden-layer activations into appropriately defined reproducing kernel Hilbert spaces and use the canonical way of constructing probability distributions over these spaces for specifying the required weight distributions in a principled way. Furthermore, we examine the practical implications of this construction for standard, finite-width networks. In particular, we derive a novel weight initialisation scheme for standard, finite-width networks that takes into account the structure of the data and information about the task at hand. We demonstrate the effectiveness of this weight initialisation approach on the MNIST, CIFAR-10 and Year Prediction MSD datasets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Implicit Acceleration and Feature Learning in Infinitely Wide Neural Networks with Bottlenecks

2021-07-01 · Etai Littwin, Omid Saremi, Shuangfei Zhai, Vimal Thilak 외

We analyze the learning dynamics of infinitely wide neural networks with a finite sized bottle-neck. Unlike the neural tangent kernel limit, a bottleneck in an otherwise infinite width network al-lows data dependent feat…

Large-width asymptotics for ReLU neural networks with $α$-Stable initializations

2022-06-16 · Stefano Favaro, Sandra Fortini, Stefano Peluchetti

There is a recent and growing literature on large-width asymptotic properties of Gaussian neural networks (NNs), namely NNs whose weights are initialized as Gaussian distributions. Two popular problems are: i) the study …

regression

Infinitely wide limits for deep Stable neural networks: sub-linear, linear and super-linear activation functions

2023-04-08 · Alberto Bordino, Stefano Favaro, Sandra Fortini

There is a growing literature on the study of large-width properties of deep Gaussian neural networks (NNs), i.e. deep NNs with Gaussian-distributed parameters or weights, and Gaussian stochastic processes. Motivated by …

On the Equivalence between Neural Network and Support Vector Machine

2021-11-11 · NeurIPS 2021 12 · Yilan Chen, Wei Huang, Lam M. Nguyen, Tsui-Wei Weng

Recent research shows that the dynamics of an infinitely wide neural network (NN) trained by gradient descent can be characterized by Neural Tangent Kernel (NTK) \citep{jacot2018neural}. Under the squared loss, the infin…

regression

Infinite-width limit of deep linear neural networks

2022-11-29 · Lénaïc Chizat, Maria Colombo, Xavier Fernández-Real, Alessio Figalli

This paper studies the infinite-width limit of deep linear neural networks initialized with random parameters. We obtain that, when the number of neurons diverges, the training dynamics converge (in a precise sense) to t…