paper-with-me

홈 › Papers

Training Infinitely Deep and Wide Transformers

2026-05-17 · Raphaël Barboni, Maarten V. de Hoop, Takashi Furuya, Gabriel Peyré arxiv

Transformers have become the dominant architecture in modern machine learning, yet the theoretical understanding of their training dynamics remains limited. This paper develops a rigorous mathematical framework for analyzing gradient-based training of transformers in the mean-field regime, where both the depth (number of layers) and width (number of attention heads) tend to infinity. While ResNet training can be understood as controlling a neural ODE, transformer training corresponds to controlling a neural PDE, due to the coupling of multiple token distributions through the attention mechanism. Our mean-field model features two types of measure representations: token distributions evolving through layers and attention parameters at each layer. We establish well-posedness of the forward pass through infinitely deep transformers, characterizing token evolution via flow maps that satisfy ODEs in function spaces. Using adjoint sensitivity analysis, we derive an explicit formula for the conditional Wasserstein gradient of the training risk, involving adjoint variables governed by backward ODEs. We prove the existence and uniqueness of gradient flow curves in the conditional Wasserstein metric space, establishing a rigorous foundation for gradient-based transformer training. A key technical contribution is providing necessary and sufficient conditions for injectivity of the Neural Tangent Kernel (NTK) for attention mechanisms: we show that NTK injectivity is equivalent to linear independence of log-sum-exp functions modulo affine functions, a condition satisfied by diverse token distributions, including discrete distributions, uniform distributions, and Gaussian mixtures. Under this NTK injectivity assumption, we prove that gradient flow converges to global minima when the initial loss is sufficiently small, eliminating spurious local minima from the optimization landscape.

📄 PDF Abstract BibTeX arXiv:2605.17660

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Implicit Acceleration and Feature Learning in Infinitely Wide Neural Networks with Bottlenecks

2021-07-01 · Etai Littwin, Omid Saremi, Shuangfei Zhai, Vimal Thilak 외

We analyze the learning dynamics of infinitely wide neural networks with a finite sized bottle-neck. Unlike the neural tangent kernel limit, a bottleneck in an otherwise infinite width network al-lows data dependent feat…

Wide and Deep Neural Networks Achieve Optimality for Classification

2022-04-29 · Adityanarayanan Radhakrishnan, Mikhail Belkin, Caroline Uhler

While neural networks are used for classification tasks across domains, a long-standing open problem in machine learning is determining whether neural networks trained using standard procedures are optimal for classifica…

Classification

Large-width asymptotics for ReLU neural networks with $α$-Stable initializations

2022-06-16 · Stefano Favaro, Sandra Fortini, Stefano Peluchetti

There is a recent and growing literature on large-width asymptotic properties of Gaussian neural networks (NNs), namely NNs whose weights are initialized as Gaussian distributions. Two popular problems are: i) the study …

regression

Infinite-width limit of deep linear neural networks

2022-11-29 · Lénaïc Chizat, Maria Colombo, Xavier Fernández-Real, Alessio Figalli

This paper studies the infinite-width limit of deep linear neural networks initialized with random parameters. We obtain that, when the number of neurons diverges, the training dynamics converge (in a precise sense) to t…

Information in Infinite Ensembles of Infinitely-Wide Neural Networks

2019-11-20 · pproximateinference AABI Symposium 2019 12 · Ravid Shwartz-Ziv, Alexander A. Alemi

In this preliminary work, we study the generalization properties of infinite ensembles of infinitely-wide neural networks. Amazingly, this model family admits tractable calculations for many information-theoretic quantit…