paper-with-me

Papers

Mamba Fusion: Learning Actions Through Questioning

2024-09-17 · Zhikang Dong, Apoorva Beedu, Jason Sheinkopf, Irfan Essa

Video Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning. While transformer-based architectures have been the de facto in vision-language training, they face challenges like quadratic computational complexity, high GPU memory usage, and difficulty with long-term dependencies. To address these limitations, we introduce MambaVL, a novel model that leverages recent advancements in selective state space modality fusion to efficiently capture long-range dependencies and learn joint representations for vision and language data. MambaVL utilizes a shared state transition matrix across both modalities, allowing the model to capture information about actions from multiple perspectives within the scene. Furthermore, we propose a question-answering task that helps guide the model toward relevant cues. These questions provide critical information about actions, objects, and environmental context, leading to enhanced performance. As a result, MambaVL achieves state-of-the-art performance in action recognition on the Epic-Kitchens-100 dataset and outperforms baseline methods in action anticipation.

📄 PDF Abstract BibTeX arXiv:2409.11513

Code (1)

dongzhikang/mambavl 공식 구현 pytorch

Tasks

Action AnticipationAction RecognitionGPUMambaQuestion Answering

Similar Papers 제목 키워드 기반

CAF-Mamba: Mamba-Based Cross-Modal Adaptive Attention Fusion for Multimodal Depression Detection

2026-01-29 · Bowen Zhou, Marc-André Fiedler, Ayoub Al-Hamadi arxiv

Depression is a prevalent mental health disorder that severely impairs daily functioning and quality of life. While recent deep learning approaches for depression detection have shown promise, most rely on limited featur…

ReMamber: Referring Image Segmentation with Mamba Twister

2024-03-26 · Yuhuan Yang, Chaofan Ma, Jiangchao Yao, Zhun Zhong 외

Referring Image Segmentation~(RIS) leveraging transformers has achieved great success on the interpretation of complex visual-language tasks. However, the quadratic computation cost makes it resource-consuming in capturi…

Image SegmentationMambaSemantic Segmentation

RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba

2024-08-16 · Andong Lu, Wanyu Wang, Chenglong Li, Jin Tang 외

Existing RGBT tracking methods often design various interaction models to perform cross-modal fusion of each layer, but can not execute the feature interactions among all layers, which plays a critical role in robust mul…

AllMambamultimodal interactionRgb-T Tracking

CAGMamba: Context-Aware Gated Cross-Modal Mamba Network for Multimodal Sentiment Analysis

2026-04-04 · Minghai Jiao, Jing Xiao, Peng Xiao, Ende Zhang 외 arxiv

Multimodal Sentiment Analysis (MSA) requires effective modeling of cross-modal interactions and contextual dependencies while remaining computationally efficient. Existing fusion approaches predominantly rely on Transfor…

Multimodal Sentiment Analysis

Bidirectional Mamba state-space model for anomalous diffusion

2024-12-10 · Maxime Lavaud, Yosef Shokeeb, Juliette Lacherez, Yacine Amarouchene 외

Characterizing anomalous diffusion is crucial in order to understand the evolution of complex stochastic systems, from molecular interactions to cellular dynamics. In this work, we characterize the performances regarding…

Mambamodel