paper-with-me

Papers

TCP-SSM: Efficient Vision State Space Models with Token-Conditioned Poles

2026-05-12 · Sara Shoouri, Morteza Tavakoli Taba, Hun-Seok Kim arxiv

State Space Models (SSMs) have emerged as a compelling alternative to attention models for long-range vision tasks, offering input-dependent recurrence with linear complexity. However, most efficient SSM variants reduce computation cost by modifying scan routes, resolutions, or traversal patterns, while largely leaving the recurrent dynamics implicit. Consequently, the model's state-dependent memory behavior is difficult to control, particularly in compact backbones where long scan paths can exceed the effective memory horizon. We propose Token-Conditioned Poles SSM (TCP-SSM), a structured selective SSM framework that improves efficiency while making recurrence dynamics explicit and interpretable through stable poles. TCP-SSM builds each scan operator with 1) real poles that model monotone or sign-alternating decay, and 2) complex-conjugate poles that capture damped oscillatory responses. Using bounded radius and angle modulation, TCP-SSM converts shared base poles into token-dependent poles, allowing each scan step to adapt its memory behavior to the current visual token while preserving pole stability. For practical scalability, we integrate grouped pole sharing with a lightweight low-rank input pathway, yielding an efficient scan operator that preserves linear-time scan complexity. Across image classification, semantic segmentation, and object detection, TCP-SSM reduces SSM computation complexity up to 44% in Vision Mamba-style models while maintaining or surpassing baseline accuracy.

📄 PDF Abstract BibTeX arXiv:2605.11563

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SegmentationImage ClassificationObject Detection

Similar Papers 제목 키워드 기반

Non-asymptotic Error Analysis of Subspace Identification for Deterministic Systems

2024-12-21 · Shuai Sun

The subspace identification method (SIM) has been extensively employed in the identification of discrete-time multiple-input multiple-output (MIMO) linear time-invariant (LTI) systems. This paper focuses on the analysis …

State Space Models

Can Graphs Help Vision SSMs See Better?

2026-05-11 · Dhruv Parikh, Anvitha Ramachandran, Haoyang Fan, Mustafa Munir 외 arxiv

Vision state space models inherit the efficiency and long-range modeling ability of Mamba-style selective scans. However, their performance depends critically on the representation of two-dimensional visual features as o…

Semantic SegmentationInstance SegmentationImage ClassificationLong-range modeling

Non-Asymptotic Analysis of Subspace Identification for Stochastic Systems Using Multiple Trajectories

2025-01-31 · Shuai Sun

This paper is concerned with the analysis of identification errors for $n$-dimensional discrete-time Linear Time-Invariant (LTI) systems with $m$ outputs and no external inputs, using Subspace Identification Methods (SIM…

Zero-shot sim-to-real transfer of tactile control policies for aggressive swing-up manipulation

2021-01-07 · Thomas Bi, Carmelo Sferrazza, Raffaello D'Andrea

This paper aims to show that robots equipped with a vision-based tactile sensor can perform dynamic manipulation tasks without prior knowledge of all the physical attributes of the objects to be manipulated. For this pur…

Finite Sample Performance Analysis of MIMO Systems Identification

2023-10-18 · Shuai Sun, Jiayun Li, Yilin Mo

This paper is concerned with the finite sample identification performance of an n dimensional discrete-time Multiple-Input Multiple-Output (MIMO) Linear Time-Invariant system, with p inputs and m outputs. We prove that t…