QueryMamba: A Mamba-Based Encoder-Decoder Architecture with a Statistical Verb-Noun Interaction Module for Video Action Forecasting @ Ego4D Long-Term Action Anticipation Challenge 2024
This report presents a novel Mamba-based encoder-decoder architecture, QueryMamba, featuring an integrated verb-noun interaction module that utilizes a statistical verb-noun co-occurrence matrix to enhance video action forecasting. This architecture not only predicts verbs and nouns likely to occur based on historical data but also considers their joint occurrence to improve forecast accuracy. The efficacy of this approach is substantiated by experimental results, with the method achieving second place in the Ego4D LTA challenge and ranking first in noun prediction accuracy.
Code (0)
등록된 구현이 없습니다.
Tasks
Action AnticipationDecoderLong Term Action AnticipationMambaSimilar Papers 제목 키워드 기반
MambaVesselNet++: A Hybrid CNN-Mamba Architecture for Medical Image Segmentation
Medical image segmentation plays an important role in computer-aided diagnosis. Traditional convolution-based U-shape segmentation architectures are usually limited by the local receptive field. Existing vision transform…
Medical Image SegmentationInstance SegmentationMultimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
Recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance but face deployment challenges due to their quadratic computational complexity, growing Key-Value cache requirements, and reliance on s…
DecoderGPUMambaState Space ModelsDM-SegNet: Dual-Mamba Architecture for 3D Medical Image Segmentation with Global Context Modeling
Accurate 3D medical image segmentation demands architectures capable of reconciling global context modeling with spatial topology preservation. While State Space Models (SSMs) like Mamba show potential for sequence model…
AnatomyBrain Tumor SegmentationDecoderImage Segmentation+7Vision Mamba-based autonomous crack segmentation on concrete, asphalt, and masonry surfaces
Convolutional neural networks (CNNs) and Transformers have shown advanced accuracy in crack detection under certain conditions. Yet, the fixed local attention can compromise the generalisation of CNNs, and the quadratic …
Crack SegmentationDecoderMambaMamba-UNet: UNet-Like Pure Visual Mamba for Medical Image Segmentation
In recent advancements in medical image analysis, Convolutional Neural Networks (CNN) and Vision Transformers (ViT) have set significant benchmarks. While the former excels in capturing local features through its convolu…
Cardiac SegmentationComputational EfficiencyDecoderImage Segmentation+5