paper-with-me

Papers

Perceptrons and localization of attention's mean-field landscape

2026-01-29 · Antonio Álvarez-López, Borjan Geshkovski, Domènec Ruiz-Balet arxiv

The forward pass of a Transformer can be seen as an interacting particle system on the unit sphere: time plays the role of layers, particles that of token embeddings, and the unit sphere idealizes layer normalization. In some weight settings the system can even be seen as a gradient flow for an explicit energy, and one can make sense of the infinite context length (mean-field) limit thanks to Wasserstein gradient flows. In this paper we study the effect of the perceptron block in this setting, and show that critical points are generically atomic and localized on subsets of the sphere.

📄 PDF Abstract BibTeX arXiv:2601.21366

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Efficient Additive Kolmogorov-Arnold Transformer for Point-Level Maize Localization in Unmanned Aerial Vehicle Imagery

2026-01-12 · Fei Li, Lang Qiao, Jiahao Fan, Yijia Xu 외 arxiv

High-resolution UAV photogrammetry has become a key technology for precision agriculture, enabling centimeter-level crop monitoring and point-level plant localization. However, point-level maize localization in UAV image…

Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape

2024-02-02 · Juno Kim, Taiji Suzuki

Large language models based on the Transformer architecture have demonstrated impressive capabilities to learn in context. However, existing theoretical studies on how this phenomenon arises are limited to the dynamics o…

In-Context Learning

Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspective

2026-05-26 · Hyunmin Cho, Woo Kyoung Han, Kyong Hwan Jin arxiv

We characterize the pre-softmax attention matrix $\mathbf{QK^\top}$ in transformers as an associative memory matrix encoding pairwise associations between input features. By decomposing this matrix into its symmetric and…

Holistically Explainable Vision Transformers

2023-01-20 · Moritz Böhle, Mario Fritz, Bernt Schiele

Transformers increasingly dominate the machine learning landscape across many tasks and domains, which increases the importance for understanding their outputs. While their attention modules provide partial insight into …

What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis

2024-10-14 · Weronika Ormaniec, Felix Dangel, Sidak Pal Singh

The Transformer architecture has inarguably revolutionized deep learning, overtaking classical architectures like multi-layer perceptrons (MLPs) and convolutional neural networks (CNNs). At its core, the attention block …