MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding
Long videos, ranging from minutes to hours, present significant challenges for current Multi-modal Large Language Models (MLLMs) due to their complex events, diverse scenes, and long-range dependencies. Direct encoding of such videos is computationally too expensive, while simple video-to-text conversion often results in redundant or fragmented content. To address these limitations, we introduce MMViR, a novel multi-modal, multi-grained structured representation for long video understanding. MMViR identifies key turning points to segment the video and constructs a three-level description that couples global narratives with fine-grained visual details. This design supports efficient query-based retrieval and generalizes well across various scenarios. Extensive evaluations across three tasks, including QA, summarization, and retrieval, show that MMViR outperforms the prior strongest method, achieving a 19.67% improvement in hour-long video understanding while reducing processing latency to 45.4% of the original.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
MM-Path: Multi-modal, Multi-granularity Path Representation Learning -- Extended Version
Developing effective path representations has become increasingly essential across various fields within intelligent transportation. Although pre-trained path representation learning models have shown improved performanc…
Representation LearningMulti-Granularity Contrastive Knowledge Distillation for Multimodal Named Entity Recognition
It is very valuable to recognize named entities from short and informal multimodal posts in this age of information explosion. Despite existing methods success in multi-modal named entity recognition (MNER), they rely on…
Knowledge DistillationMulti-modal Named Entity Recognitionnamed-entity-recognitionNamed Entity Recognition+1Multilevel Transformer For Multimodal Emotion Recognition
Multimodal emotion recognition has attracted much attention recently. Fusing multiple modalities effectively with limited labeled data is a challenging task. Considering the success of pre-trained model and fine-grained …
Emotion RecognitionMultimodal Emotion RecognitionMulti-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval
The key of the text-to-video retrieval (TVR) task lies in learning the unique similarity between each pair of text (consisting of words) and video (consisting of audio and image frames) representations. However, some pro…
RetrievalSentenceText to Video Retrievalvalid+1Learning Granularity-Unified Representations for Text-to-Image Person Re-identification
Text-to-image person re-identification (ReID) aims to search for pedestrian images of an interested identity via textual descriptions. It is challenging due to both rich intra-modal variations and significant inter-modal…
Person Re-IdentificationText based Person RetrievalText based Person Search