paper-with-me

Papers

MambaMixer: Efficient Selective State Space Models with Dual Token and Channel Selection

2024-03-29 · Ali Behrouz, Michele Santacatterina, Ramin Zabih

Recent advances in deep learning have mainly relied on Transformers due to their data dependency and ability to learn at scale. The attention module in these architectures, however, exhibits quadratic time and space in input size, limiting their scalability for long-sequence modeling. Despite recent attempts to design efficient and effective architecture backbone for multi-dimensional data, such as images and multivariate time series, existing models are either data independent, or fail to allow inter- and intra-dimension communication. Recently, State Space Models (SSMs), and more specifically Selective State Space Models, with efficient hardware-aware implementation, have shown promising potential for long sequence modeling. Motivated by the success of SSMs, we present MambaMixer, a new architecture with data-dependent weights that uses a dual selection mechanism across tokens and channels, called Selective Token and Channel Mixer. MambaMixer connects selective mixers using a weighted averaging mechanism, allowing layers to have direct access to early features. As a proof of concept, we design Vision MambaMixer (ViM2) and Time Series MambaMixer (TSM2) architectures based on the MambaMixer block and explore their performance in various vision and time series forecasting tasks. Our results underline the importance of selective mixing across both tokens and channels. In ImageNet classification, object detection, and semantic segmentation tasks, ViM2 achieves competitive performance with well-established vision models and outperforms SSM-based vision models. In time series forecasting, TSM2 achieves outstanding performance compared to state-of-the-art methods while demonstrating significantly improved computational cost. These results show that while Transformers, cross-channel attention, and MLPs are sufficient for good performance in time series forecasting, neither is necessary.

📄 PDF Abstract BibTeX arXiv:2403.19888

Code (0)

등록된 구현이 없습니다.

Tasks

channel selectionImage Classificationobject-detectionObject DetectionSemantic SegmentationState Space ModelsTime SeriesTime Series Forecasting

Similar Papers 제목 키워드 기반

NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics

2026-03-17 · Zhengzheng Tang arxiv

We ask whether a pure spiking backbone can learn large-scale language modeling from random initialization, without Transformer distillation. We introduce NeuronSpark, a 0.9B-parameter SNN language model trained with next…

SambaMixer: State of Health Prediction of Li-ion Batteries using Mamba State Space Models

2024-10-31 · José Ignacio Olalde-Verano, Sascha Kirch, Clara Pérez-Molina, Sergio Martin

The state of health (SOH) of a Li-ion battery is a critical parameter that determines the remaining capacity and the remaining lifetime of the battery. In this paper, we propose SambaMixer a novel structured state space …

Li-ion State of Health EstimationMambaState Space Models

Selective Visual Prompting in Vision Mamba

2024-12-12 · Yifeng Yao, Zichen Liu, Zhenyu Cui, Yuxin Peng 외

Pre-trained Vision Mamba (Vim) models have demonstrated exceptional performance across various computer vision tasks in a computationally efficient manner, attributed to their unique design of selective state space model…

MambaState Space ModelsVisual Prompting

Vision-Proprioception Fusion with Mamba2 in End-to-End Reinforcement Learning for Motion Control

2025-09-09 · Xiaowen Tao, Yinuo Wang, Jinzhao Zhou arxiv

End-to-end reinforcement learning (RL) for motion control trains policies directly from sensor inputs to motor commands, enabling unified controllers for different robots and tasks. However, most existing methods are eit…

Reinforcement Learning

Sessa: Selective State Space Attention

2026-04-20 · Liubomyr Horbatko arxiv

Modern sequence modeling is dominated by two families: Transformers, whose self-attention can access arbitrary elements of the visible sequence, and structured state-space models, which propagate information through an e…