paper-with-me

홈 › Papers

HAMSA: Scanning-Free Vision State Space Models via SpectralPulseNet

2026-04-16 · Badri N. Patro, Vijay S. Agneeswaran arxiv

Vision State Space Models (SSMs) like Vim, VMamba, and SiMBA rely on complex scanning strategies to adapt sequential SSMs to process 2D images, introducing computational overhead and architectural complexity. We propose HAMSA, a scanning-free SSM operating directly in the spectral domain. HAMSA introduces three key innovations: (1) simplified kernel parameterization-a single Gaussian-initialized complex kernel replacing traditional (A, B, C) matrices, eliminating discretization instabilities; (2) SpectralPulseNet (SPN)-an input-dependent frequency gating mechanism enabling adaptive spectral modulation; and (3) Spectral Adaptive Gating Unit (SAGU)-magnitude-based gating for stable gradient flow in the frequency domain. By leveraging FFT-based convolution, HAMSA eliminates sequential scanning while achieving O(L log L) complexity with superior simplicity and efficiency. On ImageNet-1K, HAMSA reaches 85.7% top-1 accuracy (state-of-the-art among SSMs), with 2.2 X faster inference than transformers (4.2ms vs 9.2ms for DeiT-S) and 1.4-1.9X speedup over scanning-based SSMs, while using less memory (2.1GB vs 3.2-4.5GB) and energy (12.5J vs 18-25J). HAMSA demonstrates strong generalization across transfer learning and dense prediction tasks.

📄 PDF Abstract BibTeX arXiv:2604.14724

Code (0)

등록된 구현이 없습니다.

Tasks

Transfer Learning

Similar Papers 제목 키워드 기반

FC-Vision: Real-Time Visibility-Aware Replanning for Occlusion-Free Aerial Target Structure Scanning in Unknown Environments

2026-02-14 · Chen Feng, Yang Xu, Shaojie Shen arxiv

Autonomous aerial scanning of target structures is crucial for practical applications, requiring online adaptation to unknown obstacles during flight. Existing methods largely emphasize collision avoidance and efficiency…

Collision Avoidance

Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics

2026-07-03 · Anvitha Ramachandran, Dhruv Parikh, Haoyang Fan, Rajgopal Kannan 외 arxiv

State Space Models (SSMs) have emerged as an alternative to Vision Transformers, yet most vision SSMs inherit directional token scanning from causal sequence modeling. While effective for sequential data, directional sca…

Instance SegmentationSemantic SegmentationObject LocalizationObject Detection

EAMamba: Efficient All-Around Vision State Space Model for Image Restoration

2025-06-27 · Yu-Cheng Lin, Yu-Syuan Xu, Hao-Wei Chen, Hsien-Kai Kuo 외

Image restoration is a key task in low-level computer vision that aims to reconstruct high-quality images from degraded inputs. The emergence of Vision Mamba, which draws inspiration from the advanced state space model M…

AllDeblurringDenoisingImage Restoration+2

XYScanNet: A State Space Model for Single Image Deblurring

2024-12-13 · Hanzhou Liu, Chengkai Liu, Jiacong Xu, Peng Jiang 외

Deep state-space models (SSMs), like recent Mamba architectures, are emerging as a promising alternative to CNN and Transformer networks. Existing Mamba-based restoration methods process visual data by leveraging a flatt…

DeblurringImage DeblurringMambaSingle Image Deblurring+1

Can Graphs Help Vision SSMs See Better?

2026-05-11 · Dhruv Parikh, Anvitha Ramachandran, Haoyang Fan, Mustafa Munir 외 arxiv

Vision state space models inherit the efficiency and long-range modeling ability of Mamba-style selective scans. However, their performance depends critically on the representation of two-dimensional visual features as o…

Semantic SegmentationInstance SegmentationImage ClassificationLong-range modeling