paper-with-me

홈 › Papers

A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models

2026-02-13 · Mugunthan Shandirasegaran, Hongkang Li, Songyang Zhang, Meng Wang, Shuai Zhang arxiv

The recent empirical success of Mamba and other selective state space models (SSMs) has renewed interest in non-attention architectures for sequence modeling, yet their theoretical foundations remain underexplored. We present a first-step analysis of generalization and learning dynamics for a simplified but representative Mamba block: a single-layer, single-head selective SSM with input-dependent gating, followed by a two-layer MLP trained via gradient descent (GD). Our study adopts a structured data model with tokens that include both class-relevant and class-irrelevant patterns under token-level noise and examines two canonical regimes: majority-voting and locality-structured data sequences. We prove that the model achieves guaranteed generalization by establishing non-asymptotic sample complexity and convergence rate bounds, which improve as the effective signal increases and the noise decreases. Furthermore, we show that the gating vector aligns with class-relevant features while ignoring irrelevant ones, thereby formalizing a feature-selection role similar to attention but realized through selective recurrence. Numerical experiments on synthetic data justify our theoretical results. Overall, our results provide principled insight into when and why Mamba-style selective SSMs learn efficiently, offering a theoretical counterpoint to Transformer-centric explanations.

📄 PDF Abstract BibTeX arXiv:2602.12499

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How Can Mamba Learn In Context with Outliers and Generalize Provably?

2025-10-01 · Hongkang Li, Songtao Lu, Xiaodong Cui, Pin-Yu Chen 외 arxiv

The Mamba model has gained significant attention for its computational advantages over Transformer-based models, while achieving comparable performance across a wide range of language tasks. Like Transformers, Mamba exhi…

Binary Classification

InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model

2026-03-08 · Youjin Wang, Jiaqiao Zhao, Rong Fu, Run Zhou 외 arxiv

Balancing fine-grained local modeling with long-range dependency capture under computational constraints remains a central challenge in sequence modeling. While Transformers provide strong token mixing, they suffer from …

Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression

2025-09-28 · Jiarui Jiang, Wei Huang, Miao Zhang, Taiji Suzuki 외 arxiv

State-space models (SSMs), particularly Mamba, emerge as an efficient Transformer alternative with linear complexity for long-sequence modeling. Recent empirical works demonstrate Mamba's in-context learning (ICL) capabi…

Block-Biased Mamba for Long-Range Sequence Processing

2025-05-13 · Annan Yu, N. Benjamin Erichson

Mamba extends earlier state space models (SSMs) by introducing input-dependent dynamics, and has demonstrated strong empirical performance across a range of domains, including language modeling, computer vision, and foun…

Inductive BiasLanguage ModelingLanguage ModellingMamba+1

KalMamba: Towards Efficient Probabilistic State Space Models for RL under Uncertainty

2024-06-21 · Philipp Becker, Niklas Freymuth, Gerhard Neumann

Probabilistic State Space Models (SSMs) are essential for Reinforcement Learning (RL) from high-dimensional, partial information as they provide concise representations for control. Yet, they lack the computational effic…

Computational EfficiencyMambaReinforcement Learning (RL)State Space Models