paper-with-me

홈 › Papers

DensePercept-NCSSD: Vision Mamba towards Real-time Dense Visual Perception with Non-Causal State Space Duality

2025-11-16 · Tushar Anand, Advik Sinha, Abhijit Das arxiv

In this work, we propose an accurate and real-time optical flow and disparity estimation model by fusing pairwise input images in the proposed non-causal selective state space for dense perception tasks. We propose a non-causal Mamba block-based model that is fast and efficient and aptly manages the constraints present in a real-time applications. Our proposed model reduces inference times while maintaining high accuracy and low GPU usage for optical flow and disparity map generation. The results and analysis, and validation in real-life scenario justify that our proposed model can be used for unified real-time and accurate 3D dense perception estimation tasks. The code, along with the models, can be found at https://github.com/vimstereo/DensePerceptNCSSD

📄 PDF Abstract BibTeX arXiv:2511.12671

Code (0)

등록된 구현이 없습니다.

Tasks

Disparity Estimation

Similar Papers 제목 키워드 기반

GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning

2024-09-22 · Huy Hoang Nguyen, An Vuong, Anh Nguyen, Ian Reid 외

Grasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with cluttered images, lengthy textual descripti…

Mamba

MambaOut: Do We Really Need Mamba for Vision?

2024-05-13 · CVPR 2025 1 · Weihao Yu, Xinchao Wang

Mamba, an architecture with RNN-like token mixer of state space model (SSM), was recently introduced to address the quadratic complexity of the attention mechanism and subsequently applied to vision tasks. Nevertheless, …

image-classificationImage ClassificationInstance SegmentationMamba+2

RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation

2024-06-06 · Jiaming Liu, Mengzhen Liu, Zhenyu Wang, Pengju An 외

A fundamental objective in robot manipulation is to enable models to comprehend visual scenes and execute actions. Although existing Vision-Language-Action (VLA) models for robots can handle a range of basic tasks, they …

Common Sense ReasoningMambaPose PredictionRobot Manipulation+2

Vision Mamba for Permeability Prediction of Porous Media

2025-10-16 · Ali Kashefi, Tapan Mukerji arxiv

Vision Mamba has recently received attention as an alternative to Vision Transformers (ViTs) for image classification. The network size of Vision Mamba scales linearly with input image resolution, whereas ViTs scale quad…

Image Classification

TrackingMiM: Efficient Mamba-in-Mamba Serialization for Real-time UAV Object Tracking

2025-07-02 · Bingxi Liu, Calvin Chen, Junhao Li, Guyang Yu 외 arxiv

The Vision Transformer (ViT) model has long struggled with the challenge of quadratic complexity, a limitation that becomes especially critical in unmanned aerial vehicle (UAV) tracking systems, where data must be proces…

Computational EfficiencyObject Tracking