paper-with-me

홈 › Papers

Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking

2025-04-11 · Huu-Loc Tran, Tinh-Anh Nguyen-Nhu, Huu-Phong Phan-Nguyen, Tien-Huy Nguyen, Nhat-Minh Nguyen-Dich, Anh Dao, Huy-Duc Do, Quan Nguyen, Hoang M. Le, Quang-Vinh Dinh

Long-form video understanding presents significant challenges for interactive retrieval systems, as conventional methods struggle to process extensive video content efficiently. Existing approaches often rely on single models, inefficient storage, unstable temporal search, and context-agnostic reranking, limiting their effectiveness. This paper presents a novel framework to enhance interactive video retrieval through four key innovations: (1) an ensemble search strategy that integrates coarse-grained (CLIP) and fine-grained (BEIT3) models to improve retrieval accuracy, (2) a storage optimization technique that reduces redundancy by selecting representative keyframes via TransNetV2 and deduplication, (3) a temporal search mechanism that localizes video segments using dual queries for start and end points, and (4) a temporal reranking approach that leverages neighboring frame context to stabilize rankings. Evaluated on known-item search and question-answering tasks, our framework demonstrates substantial improvements in retrieval precision, efficiency, and user interpretability, offering a robust solution for real-world interactive video retrieval applications.

📄 PDF Abstract BibTeX arXiv:2504.08384

Code (0)

등록된 구현이 없습니다.

Tasks

Moment RetrievalQuestion AnsweringRerankingRetrievalVideo RetrievalVideo Understanding

Similar Papers 제목 키워드 기반

UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection

2022-03-23 · CVPR 2022 1 · Ye Liu, Siyuan Li, Yang Wu, Chang Wen Chen 외

Finding relevant moments and highlights in videos according to natural language queries is a natural and highly valuable common need in the current video content explosion era. Nevertheless, jointly conducting moment ret…

DecoderHighlight DetectionMoment RetrievalNatural Language Queries+2

Unified Interactive Multimodal Moment Retrieval via Cascaded Embedding-Reranking and Temporal-Aware Score Fusion

2025-12-15 · Toan Le Ngo Thanh, Phat Ha Huu, Tan Nguyen Dang Duy, Thong Nguyen Le Minh 외 arxiv

The exponential growth of video content has created an urgent need for efficient multimodal moment retrieval systems. However, existing approaches face three critical challenges: (1) fixed-weight fusion strategies fail a…

Moment Retrieval

VeRVE: Versatile Retrieval for Videos via Unified Embeddings

2026-01-17 · Shaunak Halbe, Bhagyashree Puranik, Jayakrishnan Unnikrishnan, Kushan Thakkar 외 arxiv

Modern video retrieval systems are expected to handle diverse tasks ranging from corpus-level retrieval, fine-grained moment localization to flexible multimodal querying. Specialized architectures achieve strong retrieva…

Zero-shot Moment RetrievalZero-Shot Video Retrieval

Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval

2026-05-04 · Yiming Ding, Siyu Cao, Luyuan Jiao, Yixuan Li 외 arxiv

Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each query. This assumption does not always h…

Moment Retrieval

Who Can We Trust? Scope-Aware Video Moment Retrieval with Multi-Agent Conflict

2025-11-01 · Chaochen Wu, Guan Luo, Meiyun Zuo, Zhitao Fan arxiv

Video moment retrieval uses a text query to locate a moment from a given untrimmed video reference. Locating corresponding video moments with text queries helps people interact with videos efficiently. Current solutions …

Reinforcement LearningMoment Retrieval