paper-with-me

Papers

Implicit Acceleration and Feature Learning in Infinitely Wide Neural Networks with Bottlenecks

2021-07-01 · Etai Littwin, Omid Saremi, Shuangfei Zhai, Vimal Thilak, Hanlin Goh, Joshua M. Susskind, Greg Yang

We analyze the learning dynamics of infinitely wide neural networks with a finite sized bottle-neck. Unlike the neural tangent kernel limit, a bottleneck in an otherwise infinite width network al-lows data dependent feature learning in its bottle-neck representation. We empirically show that a single bottleneck in infinite networks dramatically accelerates training when compared to purely in-finite networks, with an improved overall performance. We discuss the acceleration phenomena by drawing similarities to infinitely wide deep linear models, where the acceleration effect of a bottleneck can be understood theoretically.

📄 PDF Abstract BibTeX arXiv:2107.00364

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fast Approximation and Estimation Bounds of Kernel Quadrature for Infinitely Wide Models

2019-02-02 · Sho Sonoda

An infinitely wide model is a weighted integration $\int \varphi(x,v) d \mu(v)$ of feature maps. This model excels at handling an infinite number of features, and thus it has been adopted to the theoretical study of deep…

Model SelectionNumerical Integration

Crystalformer: Infinitely Connected Attention for Periodic Structure Encoding

2024-03-18 · Tatsunori Taniai, Ryo Igarashi, Yuta Suzuki, Naoya Chiba 외

Predicting physical properties of materials from their crystal structures is a fundamental problem in materials science. In peripheral areas such as the prediction of molecular properties, fully connected attention netwo…

On the relationship between multivariate splines and infinitely-wide neural networks

2023-02-07 · Francis Bach

We consider multivariate splines and show that they have a random feature expansion as infinitely wide neural networks with one-hidden layer and a homogeneous activation function which is the power of the rectified linea…

On Sparsity in Overparametrised Shallow ReLU Networks

2020-06-18 · Jaume de Dios, Joan Bruna

The analysis of neural network training beyond their linearization regime remains an outstanding open question, even in the simplest setup of a single hidden-layer. The limit of infinitely wide networks provides an appea…

Open-Ended Question Answering

Neural Splines: Fitting 3D Surfaces with Infinitely-Wide Neural Networks

2020-06-24 · CVPR 2021 1 · Francis Williams, Matthew Trager, Joan Bruna, Denis Zorin

We present Neural Splines, a technique for 3D surface reconstruction that is based on random feature kernels arising from infinitely-wide shallow ReLU networks. Our method achieves state-of-the-art results, outperforming…

Surface Reconstruction