paper-with-me

Papers

A Mean-Field Analysis of Multi-Head Self-Attention under Cross-Entropy Training

2026-06-09 · Cheng Huan, Hongfwei Yuan arxiv

This paper develops a mean-field theory for a simplified single-layer causal multi-head self-attention model trained by cross-entropy minimization. Each attention head is treated as a particle in parameter space, and the empirical law of the heads is used as the large-head state variable. In the infinite-head limit, the averaged attention logits define a risk functional on probability measures, whose first variation generates a nonlinear Wasserstein gradient-flow equation. Unlike classical mean-field analyses of shallow networks that often focus on square-loss regression, the present model contains the softmax residual from the cross-entropy objective and the query-key-value structure of masked self-attention. We prove a static finite-head approximation bound for the optimal risk, characterize global minimizers through a variational support condition, and establish a quantitative finite-time propagation-of-chaos estimate comparing finite-head stochastic gradient descent with the limiting PDE. We then study the long-time behavior of the PDE: energy dissipation, convergence to the stationary set under compactness, convergence to a single stationary measure under topological or Kurdyka--Łojasiewicz assumptions, and explicit convergence rates under gradient-domination conditions. Finally, we prove local exponential stability under a Wasserstein strong-monotonicity condition and give verifiable stability and instability criteria for Dirac stationary measures. The results provide a rigorous baseline mean-field framework for attention-head training and clarify the additional compactness, landscape, and curvature assumptions needed to pass from stationarity to convergence and stability.

📄 PDF Abstract BibTeX arXiv:2606.10469

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Target-specified Sequence Labeling with Multi-head Self-attention for Target-oriented Opinion Words Extraction

2021-06-01 · NAACL 2021 4 · Yuhao Feng, Yanghui Rao, Yuyao Tang, Ninghua Wang 외

Opinion target extraction and opinion term extraction are two fundamental tasks in Aspect Based Sentiment Analysis (ABSA). Many recent works on ABSA focus on Target-oriented Opinion Words (or Terms) Extraction (TOWE), wh…

Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis (ABSA)Language ModelingLanguage Modelling+3

Homogenized Transformers

2026-04-02 · Hugo Koubbi, Borjan Geshkovski, Philippe Rigollet arxiv

We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth as a time variable, the residual stream …

Multiformer: A Head-Configurable Transformer-Based Model for Direct Speech Translation

2022-05-14 · NAACL (ACL) 2022 7 · Gerard Sant, Gerard I. Gállego, Belen Alastruey, Marta R. Costa-jussà

Transformer-based models have been achieving state-of-the-art results in several fields of Natural Language Processing. However, its direct application to speech tasks is not trivial. The nature of this sequences carries…

Translation

Low Rank Factorization for Compact Multi-Head Self-Attention

2019-11-26 · Sneha Mehta, Huzefa Rangwala, Naren Ramakrishnan

Effective representation learning from text has been an active area of research in the fields of NLP and text mining. Attention mechanisms have been at the forefront in order to learn contextual sentence representations.…

ArticlesGeneral ClassificationRepresentation LearningSentence+3

A Supervised Learning Approach For Heading Detection

2018-08-31 · Sahib Singh Budhiraja, Vijay Mago

As the Portable Document Format (PDF) file format increases in popularity, research in analysing its structure for text extraction and analysis is necessary. Detecting headings can be a crucial component of classifying a…

SensitivitySpecificity