paper-with-me

홈 › Papers

SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series

2024-03-22 · Badri N. Patro, Vijay S. Agneeswaran

Transformers have widely adopted attention networks for sequence mixing and MLPs for channel mixing, playing a pivotal role in achieving breakthroughs across domains. However, recent literature highlights issues with attention networks, including low inductive bias and quadratic complexity concerning input sequence length. State Space Models (SSMs) like S4 and others (Hippo, Global Convolutions, liquid S4, LRU, Mega, and Mamba), have emerged to address the above issues to help handle longer sequence lengths. Mamba, while being the state-of-the-art SSM, has a stability issue when scaled to large networks for computer vision datasets. We propose SiMBA, a new architecture that introduces Einstein FFT (EinFFT) for channel modeling by specific eigenvalue computations and uses the Mamba block for sequence modeling. Extensive performance studies across image and time-series benchmarks demonstrate that SiMBA outperforms existing SSMs, bridging the performance gap with state-of-the-art transformers. Notably, SiMBA establishes itself as the new state-of-the-art SSM on ImageNet and transfer learning benchmarks such as Stanford Car and Flower as well as task learning benchmarks as well as seven time series benchmark datasets. The project page is available on this website ~\url{https://github.com/badripatro/Simba}.

📄 PDF Abstract BibTeX arXiv:2403.15360

Code (3)

badripatro/simba 공식 구현 pytorch
MindSpore-scientific-2/code-10/tree/main/Simba mindspore
MindSpore-scientific-2/code-11/tree/main/Simba mindspore

Tasks

Inductive BiasMambaState Space ModelsTime SeriesTransfer Learning

Similar Papers 제목 키워드 기반

SimBase: A Simple Baseline for Temporal Video Grounding

2024-11-12 · Peijun Bao, Alex C. Kot

This paper presents SimBase, a simple yet effective baseline for temporal video grounding. While recent advances in temporal grounding have led to impressive performance, they have also driven network architectures towar…

Video Grounding

Simba: Mamba augmented U-ShiftGCN for Skeletal Action Recognition in Videos

2024-04-11 · Soumyabrata Chaudhuri, Saumik Bhattacharya

Skeleton Action Recognition (SAR) involves identifying human actions using skeletal joint coordinates and their interconnections. While plain Transformers have been attempted for this task, they still fall short compared…

Action RecognitionAction Recognition In VideosMamba

Exploring State-Space-Model based Language Model in Music Generation

2025-07-09 · Wei-Jaw Lee, Fang-Chih Hsieh, Xuanjun Chen, Fang-Duo Tsai 외 arxiv

The recent surge in State Space Models (SSMs), particularly the emergence of Mamba, has established them as strong alternatives or complementary modules to Transformers across diverse domains. In this work, we aim to exp…

Text-to-Music Generation

Sparsified State-Space Models are Efficient Highway Networks

2025-05-27 · Woomin Song, Jihoon Tack, Sangwoo Mo, Seunghyuk Oh 외

State-space models (SSMs) offer a promising architecture for sequence modeling, providing an alternative to Transformers by replacing expensive self-attention with linear recurrences. In this paper, we propose a simple y…

MambaState Space Models

HAMSA: Scanning-Free Vision State Space Models via SpectralPulseNet

2026-04-16 · Badri N. Patro, Vijay S. Agneeswaran arxiv

Vision State Space Models (SSMs) like Vim, VMamba, and SiMBA rely on complex scanning strategies to adapt sequential SSMs to process 2D images, introducing computational overhead and architectural complexity. We propose …

Transfer Learning