paper-with-me

홈 › Papers

Infinite Limits of Multi-head Transformer Dynamics

2024-05-24 · Blake Bordelon, Hamza Tahir Chaudhry, Cengiz Pehlevan

In this work, we analyze various scaling limits of the training dynamics of transformer models in the feature learning regime. We identify the set of parameterizations that admit well-defined infinite width and depth limits, allowing the attention layers to update throughout training--a relevant notion of feature learning in these models. We then use tools from dynamical mean field theory (DMFT) to analyze various infinite limits (infinite key/query dimension, infinite heads, and infinite depth) which have different statistical descriptions depending on which infinite limit is taken and how attention layers are scaled. We provide numerical evidence of convergence to the limits and discuss how the parameterization qualitatively influences learned features.

📄 PDF Abstract BibTeX arXiv:2405.15712

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Gaussian Process Limit Reveals Structural Benefits of Graph Transformers

2026-03-18 · Nil Ayday, Lingchu Yang, Debarghya Ghoshdastidar arxiv

Graph transformers are the state-of-the-art for learning from graph-structured data and are empirically known to avoid several pitfalls of message-passing architectures. However, there is limited theoretical analysis on …

Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks

2023-10-03 · Greg Yang, Dingli Yu, Chen Zhu, Soufiane Hayou

By classifying infinite-width neural networks and identifying the *optimal* limit, Tensor Programs IV and V demonstrated a universal way, called $\mu$P, for *widthwise hyperparameter transfer*, i.e., predicting optimal h…

Diversity

Self-Attention And Beyond the Infinite: Towards Linear Transformers with Infinite Self-Attention

2026-02-26 · Giorgio Roffo, Hazem Abdelkawy, Nilli Lavie, Luke Palmer arxiv

The quadratic cost of softmax attention limits Transformer scalability in high-resolution vision. We introduce Infinite Self-Attention (InfSA), a spectral reformulation that treats each attention layer as a diffusion ste…

Uniform Scaling Limits in AdamW-Trained Transformers

2026-05-11 · William Gibson, Christoph Reisinger arxiv

We study the large-depth limit of transformers trained with AdamW, by modelling the hidden-state dynamics as an interacting particle system (IPS) coupled through the attention mechanism. Under appropriate scaling of the …

Infinite Neural Network Quantum States: Entanglement and Training Dynamics

2021-12-01 · Di Luo, James Halverson

We study infinite limits of neural network quantum states ($\infty$-NNQS), which exhibit representation power through ensemble statistics, and also tractable gradient descent dynamics. Ensemble averages of Renyi entropie…