paper-with-me

Papers

MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding

2026-01-09 · Zizhong Li, Haopeng Zhang, Jiawei Zhang arxiv

Long videos, ranging from minutes to hours, present significant challenges for current Multi-modal Large Language Models (MLLMs) due to their complex events, diverse scenes, and long-range dependencies. Direct encoding of such videos is computationally too expensive, while simple video-to-text conversion often results in redundant or fragmented content. To address these limitations, we introduce MMViR, a novel multi-modal, multi-grained structured representation for long video understanding. MMViR identifies key turning points to segment the video and constructs a three-level description that couples global narratives with fine-grained visual details. This design supports efficient query-based retrieval and generalizes well across various scenarios. Extensive evaluations across three tasks, including QA, summarization, and retrieval, show that MMViR outperforms the prior strongest method, achieving a 19.67% improvement in hour-long video understanding while reducing processing latency to 45.4% of the original.

📄 PDF Abstract BibTeX arXiv:2601.05495

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MM-Path: Multi-modal, Multi-granularity Path Representation Learning -- Extended Version

2024-11-27 · Ronghui Xu, Hanyin Cheng, Chenjuan Guo, Hongfan Gao 외

Developing effective path representations has become increasingly essential across various fields within intelligent transportation. Although pre-trained path representation learning models have shown improved performanc…

Representation Learning

Multi-Granularity Contrastive Knowledge Distillation for Multimodal Named Entity Recognition

2021-11-16 · ACL ARR November 2021 11 · Anonymous

It is very valuable to recognize named entities from short and informal multimodal posts in this age of information explosion. Despite existing methods success in multi-modal named entity recognition (MNER), they rely on…

Knowledge DistillationMulti-modal Named Entity Recognitionnamed-entity-recognitionNamed Entity Recognition+1

Multilevel Transformer For Multimodal Emotion Recognition

2022-10-26 · Junyi He, Meimei Wu, Meng Li, Xiaobo Zhu 외

Multimodal emotion recognition has attracted much attention recently. Fusing multiple modalities effectively with limited labeled data is a challenging task. Considering the success of pre-trained model and fine-grained …

Emotion RecognitionMultimodal Emotion Recognition

Multi-Granularity and Multi-modal Feature Interaction Approach for Text Video Retrieval

2024-06-21 · Wenjun Li, Shudong Wang, Dong Zhao, Shenghui Xu 외

The key of the text-to-video retrieval (TVR) task lies in learning the unique similarity between each pair of text (consisting of words) and video (consisting of audio and image frames) representations. However, some pro…

RetrievalSentenceText to Video Retrievalvalid+1

Learning Granularity-Unified Representations for Text-to-Image Person Re-identification

2022-07-16 · Zhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin 외

Text-to-image person re-identification (ReID) aims to search for pedestrian images of an interested identity via textual descriptions. It is challenging due to both rich intra-modal variations and significant inter-modal…

Person Re-IdentificationText based Person RetrievalText based Person Search