paper-with-me

홈 › Papers

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

2026-06-18 · Jindi Lv, Aoyu Li, Yuhao Zhou, Zheng Zhu, Xiaofeng Wang, Qing Ye, Yueqi Duan, Wentao Feng, Jiancheng Lv arxiv

Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced Mamba variants, these models exhibit a severe performance collapse. We attribute this degradation to the spatially agnostic nature of existing reduction methods, which violate the two-dimensional structural premise required by the selective scanning mechanism. In this work, we propose STORM, a spatial-aware token reduction framework designed to maintain structural integrity throughout the compression process. STORM reformulates reduction into a structured operation on spatial units, enforcing localized constraints to maintain both grid topology and neighborhood coherence. As a plug-and-play module, STORM equips existing reduction pipelines with explicit spatial awareness without any training. Empirical results demonstrate that STORM achieves state-of-the-art pruning accuracy across diverse vision Mamba backbones under training-free settings. Notably, STORM delivers a substantial accuracy recovery on VMamba, outperforming prior methods by up to 63.3\% in top-1 accuracy. Meanwhile, STORM incurs only a 1.0\% accuracy drop on PlainMamba, achieving performance comparable to ViT.

📄 PDF Abstract BibTeX arXiv:2606.19932

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Text-Vision Co-Instructed Image Editing

2026-06-15 · Chenxi Xie, Yuhui Wu, Qiaosi Yi, Lei Zhang arxiv

Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are limited by the coarse granularity of spat…

Image ManipulationImage Editing

Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing

2026-01-25 · Weiyu Zhang, Yuan Hu, Yong Li, Yu Liu arxiv

Unified remote sensing multimodal models exhibit a pronounced spatial reversal curse: Although they can accurately recognize and describe object locations in images, they often fail to faithfully execute the same spatial…

Text-to-Image GenerationImage CaptioningVisual Grounding

Towards Spatially-Aware and Optimally Faithful Concept-Based Explanations

2025-04-15 · Shubham Kumar, Dwip Dalal, Narendra Ahuja

Post-hoc, unsupervised concept-based explanation methods (U-CBEMs) are a promising tool for generating semantic explanations of the decision-making processes in deep neural networks, having applications in both model imp…

Decision Making

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning

2026-04-19 · Yian Li, Yang Jiao, Bin Zhu, Tianwen Qian 외 arxiv

Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, r…

multimodal generationSpatial Reasoning

Trace, Verify, and Correct: A Training-Free Framework for Spatial Reasoning in Multimodal LLMs

2026-08-05 · Yang Yang, Jiawei Chen, Tairan Chen, Zhaoxia Yin arxiv

Although Multimodal Large Language Models (MLLMs) have made substantial progress, their spatial reasoning may still produce intermediate judgments inconsistent with the input image, allowing errors to propagate through t…

Spatial Reasoning