paper-with-me

홈 › Papers

Gaussian Mixture Attention: Linear-Time Sequence Mixing via Probabilistic Latent Routing

2026-06-09 · Yongchao Huang, Hassan Raza arxiv

The dense token-to-token interaction pattern of standard dot-product attention remains a central bottleneck in scaling Transformer architectures to long contexts. We introduce \textbf{Gaussian Mixture Attention (GMA)}, a probabilistic attention-style sequence mixer that replaces explicit pairwise query--key comparison with routing through $K$ learned Gaussian mixture components. Queries and keys are mapped to posterior \textit{responsibility} vectors over a shared latent routing space; their overlap defines an implicit responsibility-space affinity, while values are written into and read from a $K$-slot latent memory. By exploiting the associativity of matrix multiplication, GMA avoids materializing the induced $N\times N$ affinity matrix and instead uses two responsibility matrices whose dominant activation storage scales as $\mathcal{O}(NK)$ rather than $\mathcal{O}(N^2)$ for fixed $K$. We formulate bidirectional and causal variants of GMA, provide an end-to-end differentiable parameterization of the Gaussian mixture components, and analyze its responsibility-modulated gradient structure, constrained non-negative low-rank affinity interpretation, and local routing stability. Empirically, GMA exhibits the intended fixed-$K$ linear memory scaling and is competitive with attention-style baselines on long-context classification, while causal GMA improves over tested linear/random-feature attention variants on WikiText-103 but remains behind optimized causal SDPA and Mamba in the current implementation. Analysis of learned responsibilities further shows broad component usage and moderate alignment with surface-form token categories, supporting GMA as a probabilistic, interpretable, fixed-$K$ linear-time attention-style alternative rather than a universal replacement for optimized softmax attention or state-space models.

📄 PDF Abstract BibTeX arXiv:2606.18283

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FRMDN: Flow-based Recurrent Mixture Density Network

2020-08-05 · Seyedeh Fatemeh Razavi, Reshad Hosseini, Tina Behzad

The class of recurrent mixture density networks is an important class of probabilistic models used extensively in sequence modeling and sequence-to-sequence mapping applications. In this class of models, the density of a…

Online Vector Quantized Attention

2026-02-03 · Nick Alonso, Tomas Figliolia, Beren Millidge arxiv

Standard sequence mixing layers used in language models struggle to balance efficiency and performance. Self-attention performs well on long context tasks but has expensive quadratic compute and linear memory costs, whil…

Improving Transformers with Probabilistic Attention Keys

2021-10-16 · Tam Nguyen, Tan M. Nguyen, Dung D. Le, Duy Khuong Nguyen 외

Multi-head attention is a driving force behind state-of-the-art transformers, which achieve remarkable performance across a variety of natural language processing (NLP) and computer vision tasks. It has been observed tha…

Language ModelingLanguage Modelling

Rethinking Nonlinearity: Trainable Gaussian Mixture Modules for Modern Neural Architectures

2025-10-08 · Weiguo Lu, Gangnan Yuan, Hong-kun Zhang, Shangyang Li arxiv

Neural networks in general, from MLPs and CNNs to attention-based Transformers, are constructed from layers of linear combinations followed by nonlinear operations such as ReLU, Sigmoid, or Softmax. Despite their strengt…

Density Steering of Gaussian Mixture Models for Discrete-Time Linear Systems

2023-11-14 · Isin M. Balci, Efstathios Bakolas

In this paper, we study the finite-horizon optimal density steering problem for discrete-time stochastic linear dynamical systems. Specifically, we focus on steering probability densities represented as Gaussian mixture …