paper-with-me

홈 › Papers

Mamba Modulation: On the Length Generalization of Mamba

2025-09-23 · Peng Lu, Jerry Huang, Qiuhao Zeng, Xinyu Wang, Boxing Chen, Philippe Langlais, Yufei Cui arxiv

The quadratic complexity of the attention mechanism in Transformer models has motivated the development of alternative architectures with sub-quadratic scaling, such as state-space models. Among these, Mamba has emerged as a leading architecture, achieving state-of-the-art results across a range of language modeling tasks. However, Mamba's performance significantly deteriorates when applied to contexts longer than those seen during pre-training, revealing a sharp sensitivity to context length extension. Through detailed analysis, we attribute this limitation to the out-of-distribution behaviour of its state-space dynamics, particularly within the parameterization of the state transition matrix $\mathbf{A}$. Unlike recent works which attribute this sensitivity to the vanished accumulation of discretization time steps, $\exp(-\sum_{t=1}^NΔ_t)$, we establish a connection between state convergence behavior as the input length approaches infinity and the spectrum of the transition matrix $\mathbf{A}$, offering a well-founded explanation of its role in length extension. Next, to overcome this challenge, we propose an approach that applies spectrum scaling to pre-trained Mamba models to enable robust long-context generalization by selectively modulating the spectrum of $\mathbf{A}$ matrices in each layer. We show that this can significantly improve performance in settings where simply modulating $Δ_t$ fails, validating our insights and providing avenues for better length generalization of state-space models with structured transition matrices.

📄 PDF Abstract BibTeX arXiv:2509.19633

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Probing Length Generalization in Mamba via Image Reconstruction

2026-03-12 · Jan Rathjens, Robin Schiewer, Laurenz Wiskott, Anand Subramoney arxiv

Mamba has attracted widespread interest as a general-purpose sequence model due to its low computational complexity and competitive performance relative to transformers. However, its performance can degrade when inferenc…

Image Reconstruction

DeciMamba: Exploring the Length Extrapolation Potential of Mamba

2024-06-20 · Assaf Ben-Kish, Itamar Zimerman, Shady Abu-Hussein, Nadav Cohen 외

Long-range sequence processing poses a significant challenge for Transformers due to their quadratic complexity in input length. A promising alternative is Mamba, which demonstrates high performance and achieves Transfor…

Mamba

Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative

2025-08-12 · Xi Xuan, Zimo Zhu, Wenxin Zhang, Yi-Cheng Lin 외 arxiv

Advances in speech synthesis intensify security threats, motivating real-time deepfake detection research. We investigate whether bidirectional Mamba can serve as a competitive alternative to Self-Attention in detecting …

DeepFake DetectionSpeech Synthesis

IRSRMamba: Infrared Image Super-Resolution via Mamba-based Wavelet Transform Feature Modulation Model

2024-05-16 · Yongsong Huang, Tomo Miyazaki, Xiaofeng Liu, Shinichiro Omachi

Infrared image super-resolution demands long-range dependency modeling and multi-scale feature extraction to address challenges such as homogeneous backgrounds, weak edges, and sparse textures. While Mamba-based state-sp…

Image EnhancementImage ReconstructionImage Super-ResolutionInfrared image super-resolution+4

TransMamba: Flexibly Switching between Transformer and Mamba

2025-03-31 · Yixing Li, Ruobing Xie, Zhen Yang, Xingwu Sun 외

Transformers are the cornerstone of modern large language models, but their quadratic computational complexity limits efficiency in long-sequence processing. Recent advancements in Mamba, a state space model (SSM) with l…

MambaScheduling