paper-with-me

홈 › Papers

Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection

2025-01-18 · Yifang Xu, Yunzhuo Sun, Benxiang Zhai, Zien Xie, Youyao Jia, Sidan Du

Given a video and a linguistic query, video moment retrieval and highlight detection (MR&HD) aim to locate all the relevant spans while simultaneously predicting saliency scores. Most existing methods utilize RGB images as input, overlooking the inherent multi-modal visual signals like optical flow and depth. In this paper, we propose a Multi-modal Fusion and Query Refinement Network (MRNet) to learn complementary information from multi-modal cues. Specifically, we design a multi-modal fusion module to dynamically combine RGB, optical flow, and depth map. Furthermore, to simulate human understanding of sentences, we introduce a query refinement module that merges text at different granularities, containing word-, phrase-, and sentence-wise levels. Comprehensive experiments on QVHighlights and Charades datasets indicate that MRNet outperforms current state-of-the-art methods, achieving notable improvements in MR-mAP@Avg (+3.41) and HD-HIT@1 (+3.46) on QVHighlights.

📄 PDF Abstract BibTeX arXiv:2501.10692

Code (0)

등록된 구현이 없습니다.

Tasks

AvgHighlight DetectionMoment RetrievalOptical Flow EstimationRetrievalSentence

Similar Papers 제목 키워드 기반

CONQUER: Contextual Query-aware Ranking for Video Corpus Moment Retrieval

2021-09-21 · Zhijian Hou, Chong-Wah Ngo, Wing Kwong Chan

This paper tackles a recently proposed Video Corpus Moment Retrieval task. This task is essential because advanced video retrieval applications should enable users to retrieve a precise moment from a large video corpus. …

Corpus Video Moment RetrievalMoment Retrievalorpus Video Moment RetrievalRepresentation Learning+4

VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval

2024-12-02 · Dhiman Paul, Md Rizwan Parvez, Nabeel Mohammed, Shafin Rahman

Video Highlight Detection and Moment Retrieval (HD/MR) are essential in video analysis. Recent joint prediction transformer models often overlook their cross-task dynamics and video-text alignment and refinement. Moreove…

Highlight DetectionMoment Retrieval

Towards Visual Query Localization in the 3D World

2026-05-02 · Liang Peng, Bohan Tan, Zhipeng Zhang, Haobo Li 외 arxiv

Visual query localization (VQL) aims to predict the spatio-temporal response of the most recent occurrence in a sequence given a query. Currently, most research focuses on visual query localization in 2D videos, while it…

Point Clouds

Dual Prototype Attention for Unsupervised Video Object Segmentation

2022-11-22 · CVPR 2024 1 · Suhwan Cho, Minhyeok Lee, Seunghoon Lee, Dogyoon Lee 외

Unsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; an…

ObjectSemantic SegmentationUnsupervised Video Object SegmentationVideo Object Segmentation+1

Cross-Modal Interaction Networks for Query-Based Moment Retrieval in Videos

2019-06-06 · Zhu Zhang, Zhijie Lin, Zhou Zhao, Zhenxin Xiao

Query-based moment retrieval aims to localize the most relevant moment in an untrimmed video according to the given natural language query. Existing works often only focus on one aspect of this emerging task, such as the…

Moment RetrievalNatural Language QueriesRepresentation LearningRetrieval