paper-with-me

홈 › Papers

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving

2025-01-08 · Siran Chen, Yuxiao Luo, Yue Ma, Yu Qiao, Yali Wang

With the prevalence of Multimodal Large Language Models(MLLMs), autonomous driving has encountered new opportunities and challenges. In particular, multi-modal video understanding is critical to interactively analyze what will happen in the procedure of autonomous driving. However, videos in such a dynamical scene that often contains complex spatial-temporal movements, which restricts the generalization capacity of the existing MLLMs in this field. To bridge the gap, we propose a novel Hierarchical Mamba Adaptation (H-MBA) framework to fit the complicated motion changes in autonomous driving videos. Specifically, our H-MBA consists of two distinct modules, including Context Mamba (C-Mamba) and Query Mamba (Q-Mamba). First, C-Mamba contains various types of structure state space models, which can effectively capture multi-granularity video context for different temporal resolutions. Second, Q-Mamba flexibly transforms the current frame as the learnable query, and attentively selects multi-granularity video context into query. Consequently, it can adaptively integrate all the video contexts of multi-scale temporal resolutions to enhance video understanding. Via a plug-and-play paradigm in MLLMs, our H-MBA shows the remarkable performance on multi-modal video tasks in autonomous driving, e.g., for risk object detection, it outperforms the previous SOTA method with 5.5% mIoU improvement.

📄 PDF Abstract BibTeX arXiv:2501.04302

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingMambaobject-detectionObject DetectionState Space ModelsVideo Understanding

Methods 이 논문이 사용한 방법론

Mamba Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module.…

Similar Papers 제목 키워드 기반

ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning

2026-04-09 · Daichi Yashima, Shuhei Kurita, Yusuke Oda, Shuntaro Suzuki 외 arxiv

In this study, we focus on video captioning by fully open multimodal large language models (MLLMs). The comprehension of visual sequences is challenging because of their intricate temporal dependencies and substantial se…

Video Captioning

DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection

2024-09-24 · Jiaxin Ye, Junping Zhang, Hongming Shan

Depression is a common mental disorder that affects millions of people worldwide. Although promising, current multimodal methods hinge on aligned or aggregated multimodal fusion, suffering two significant limitations: (i…

Depression DetectionMamba

Samba+: General and Accurate Salient Object Detection via A More Unified Mamba-based Framework

2026-02-02 · Wenzhuo Zhao, Keren Fu, Jiahao He, Xiaohong Liu 외 arxiv

Existing salient object detection (SOD) models are generally constrained by the limited receptive fields of convolutional neural networks (CNNs) and quadratic computational complexity of Transformers. Recently, the emerg…

Computational EfficiencySalient Object DetectionContinual Learning

HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling

2025-10-27 · Joungbin An, Kristen Grauman arxiv

Video temporal grounding, the task of localizing the start and end times of a natural language query in untrimmed video, requires capturing both global context and fine-grained temporal detail. This challenge is particul…

SurvMamba: State Space Model with Multi-grained Multi-modal Interaction for Survival Prediction

2024-04-11 · Ying Chen, Jiajing Xie, Yuxiang Lin, Yuhang Song 외

Multi-modal learning that combines pathological images with genomic data has significantly enhanced the accuracy of survival prediction. Nevertheless, existing methods have not fully utilized the inherent hierarchical st…

MambaPredictionSurvival Predictionwhole slide images