paper-with-me

홈 › Papers

Dynamic Vision Mamba

2025-04-07 · Mengxuan Wu, Zekai Li, Zhiyuan Liang, Moyang Li, Xuanlei Zhao, Samir Khaki, Zheng Zhu, Xiaojiang Peng, Konstantinos N. Plataniotis, Kai Wang, Wangbo Zhao, Yang You

Mamba-based vision models have gained extensive attention as a result of being computationally more efficient than attention-based models. However, spatial redundancy still exists in these models, represented by token and block redundancy. For token redundancy, we analytically find that early token pruning methods will result in inconsistency between training and inference or introduce extra computation for inference. Therefore, we customize token pruning to fit the Mamba structure by rearranging the pruned sequence before feeding it into the next Mamba block. For block redundancy, we allow each image to select SSM blocks dynamically based on an empirical observation that the inference speed of Mamba-based vision models is largely affected by the number of SSM blocks. Our proposed method, Dynamic Vision Mamba (DyVM), effectively reduces FLOPs with minor performance drops. We achieve a reduction of 35.2\% FLOPs with only a loss of accuracy of 1.7\% on Vim-S. It also generalizes well across different Mamba vision model architectures and different vision tasks. Our code will be made public.

📄 PDF Abstract BibTeX arXiv:2504.04787

Code (1)

nus-hpc-ai-lab/dyvm 공식 구현 pytorch

Tasks

Mamba

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Pruning 설명 없음
Mamba Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.…

Similar Papers 제목 키워드 기반

FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba

2024-04-15 · Xinyu Xie, Yawen Cui, Tao Tan, Xubin Zheng 외

Multimodal image fusion aims to integrate information from different imaging techniques to produce a comprehensive, detail-rich single image for downstream vision tasks. Existing methods based on local convolutional neur…

Infrared And Visible Image FusionMambaState Space Models

OuroMamba: A Data-Free Quantization Framework for Vision Mamba Models

2025-03-13 · Akshat Ramachandran, Mingyu Lee, Huan Xu, Souvik Kundu 외

We present OuroMamba, the first data-free post-training quantization (DFQ) method for vision Mamba-based models (VMMs). We identify two key challenges in enabling DFQ for VMMs, (1) VMM's recurrent state transitions restr…

channel selectionContrastive LearningData Free QuantizationGPU+3

MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba

2025-11-29 · Shanhui Liu, Rui Xu, Yunke Wang arxiv

Vision Mamba has emerged as a promising and efficient alternative to Vision Transformers, yet its efficiency remains fundamentally constrained by the number of input tokens. Existing token reduction approaches typically …

DAMamba: Vision State Space Model with Dynamic Adaptive Scan

2025-02-18 · Tanzhe Li, Caoshuo Li, Jiayi Lyu, Hongjuan Pei 외

State space models (SSMs) have recently garnered significant attention in computer vision. However, due to the unique characteristics of image data, adapting SSMs from natural language processing to computer vision has n…

image-classificationImage ClassificationInstance SegmentationMamba+4

MambaEVT: Event Stream based Visual Object Tracking using State Space Model

2024-08-20 · Xiao Wang, Chao Wang, Shiao Wang, Xixi Wang 외

Event camera-based visual tracking has drawn more and more attention in recent years due to the unique imaging principle and advantages of low energy consumption, high dynamic range, and dense temporal resolution. Curren…

MambaObject LocalizationObject TrackingVisual Object Tracking+1