paper-with-me

홈 › Papers

How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models

2025-12-17 · Ali Ghodsi arxiv

Sequence modeling has produced diverse architectures -- from classical recurrent neural networks to modern Transformers and state space models (SSMs) -- yet a unified theoretical understanding of expressivity and trainability trade-offs remains limited. We introduce a unified framework that represents a broad class of sequence maps via an input-dependent effective interaction operator $W_{ij}(X)$, making explicit two recurring construction patterns: (i) the Unified Factorized Framework (Explicit) (attention-style mixing), in which $W_{ij}(X)$ varies through scalar coefficients applied to shared value maps, and (ii) Structured Dynamics (Implicit) (state-space recurrences), in which $W_{ij}$ is induced by a latent dynamical system. Using this framework, we derive three theoretical results. First, we establish the Interaction Rank Gap: models in the Unified Factorized Framework, such as single-head attention, are constrained to a low-dimensional operator span and cannot represent certain structured dynamical maps. Second, we prove an Equivalence (Head-Count) Theorem showing that, within our multi-head factorized class, representing a linear SSM whose lag operators span a $k$-dimensional subspace on length-$n$ sequences requires and is achievable with $H=k$ heads. Third, we prove a Gradient Highway Result, showing that attention layers admit inputs with distance-independent gradient paths, whereas stable linear dynamics exhibit distance-dependent gradient attenuation. Together, these results formalize a fundamental trade-off between algebraic expressivity (interaction/operator span) and long-range gradient propagation, providing theoretical grounding for modern sequence architecture design.

📄 PDF Abstract BibTeX arXiv:2512.15115

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unifying Attention Heads and Task Vectors via Hidden State Geometry in In-Context Learning

2025-05-24 · Haolin Yang, Hakaze Cho, Yiqiao Zhong, Naoya Inoue

The unusual properties of in-context learning (ICL) have prompted investigations into the internal mechanisms of large language models. Prior work typically focuses on either special attention heads or task vectors at sp…

In-Context Learning

Are Sixteen Heads Really Better than One?

2019-05-25 · NeurIPS 2019 12 · Paul Michel, Omer Levy, Graham Neubig

Attention is a powerful and ubiquitous mechanism for allowing neural models to focus on particular salient pieces of information by taking their weighted average when making predictions. In particular, multi-headed atten…

Localizing Task Recognition and Task Learning in In-Context Learning via Attention Head Analysis

2025-09-29 · Haolin Yang, Hakaze Cho, Naoya Inoue arxiv

We investigate the mechanistic underpinnings of in-context learning (ICL) in large language models by reconciling two dominant perspectives: the component-level analysis of attention heads and the holistic decomposition …

A Language Model with Limited Memory Capacity Captures Interference in Human Sentence Processing

2023-10-24 · William Timkey, Tal Linzen

Two of the central factors believed to underpin human sentence processing difficulty are expectations and retrieval from working memory. A recent attempt to create a unified cognitive model integrating these two factors …

Language ModelingLanguage ModellingRetrievalSentence

HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation

2026-07-22 · Jinliang Shen, Lianghao Su, Zheming Li, Kang He 외 arxiv

Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attention the dominant inference cost, especiall…

Video Generation