paper-with-me

Papers

Proportional infinite-width infinite-depth limit for deep linear neural networks

2024-11-22 · Federico Bassetti, Lucia Ladelli, Pietro Rotondo

We study the distributional properties of linear neural networks with random parameters in the context of large networks, where the number of layers diverges in proportion to the number of neurons per layer. Prior works have shown that in the infinite-width regime, where the number of neurons per layer grows to infinity while the depth remains fixed, neural networks converge to a Gaussian process, known as the Neural Network Gaussian Process. However, this Gaussian limit sacrifices descriptive power, as it lacks the ability to learn dependent features and produce output correlations that reflect observed labels. Motivated by these limitations, we explore the joint proportional limit in which both depth and width diverge but maintain a constant ratio, yielding a non-Gaussian distribution that retains correlations between outputs. Our contribution extends previous works by rigorously characterizing, for linear activation functions, the limiting distribution as a nontrivial mixture of Gaussians.

📄 PDF Abstract BibTeX arXiv:2411.15267

Code (0)

등록된 구현이 없습니다.

Tasks

Descriptive

Methods 이 논문이 사용한 방법론

Gaussian Process Gaussian Processes are non-parametric models for approximating functions. They rely upon a measure of similarity between points (the kernel function) to predict the value for…

Similar Papers 제목 키워드 기반

On the infinite-depth limit of finite-width neural networks

2022-10-03 · Soufiane Hayou

In this paper, we study the infinite-depth limit of finite-width residual neural networks with random Gaussian weights. With proper scaling, we show that by fixing the width and taking the depth to infinity, the pre-acti…

The Shaped Transformer: Attention Models in the Infinite Depth-and-Width Limit

2023-06-30 · NeurIPS 2023 11 · Lorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He 외

In deep learning theory, the covariance matrix of the representations serves as a proxy to examine the network's trainability. Motivated by the success of Transformers, we study the covariance matrix of a modified Softma…

Deep AttentionLearning Theory

The Future is Log-Gaussian: ResNets and Their Infinite-Depth-and-Width Limit at Initialization

2021-06-07 · NeurIPS 2021 12 · Mufan Bill Li, Mihai Nica, Daniel M. Roy

Theoretical results show that neural networks can be approximated by Gaussian processes in the infinite-width limit. However, for fully connected networks, it has been previously shown that for any fixed network width, $…

Gaussian Processes

How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences

2026-05-06 · Mariia Seleznova arxiv

We study signal propagation in linear recurrent models at finite width. While existing signal propagation theory relies predominantly on the infinite-width limit, it remains unclear for how long that approximation remain…

Infinite Limits of Multi-head Transformer Dynamics

2024-05-24 · Blake Bordelon, Hamza Tahir Chaudhry, Cengiz Pehlevan

In this work, we analyze various scaling limits of the training dynamics of transformer models in the feature learning regime. We identify the set of parameterizations that admit well-defined infinite width and depth lim…