paper-with-me

홈 › Papers

Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval

2026-05-04 · Yiming Ding, Siyu Cao, Luyuan Jiao, Yixuan Li, Zitong Wang, Zhiyong Liu, Lu Zhang arxiv

Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each query. This assumption does not always hold in real-world scenarios, where queries may correspond to multiple or no moments. Thus, we formulate Generalized Moment Retrieval (GMR), a unified setting that requires retrieving the complete set of relevant moments or predicting an empty set. To enable systematic study of GMR, we introduce Soccer-GMR, a large-scale benchmark built on challenging soccer videos that reflect general GMR scenarios, with realistic negative and positive queries. The benchmark is constructed via a duration-flexible semi-automated pipeline with human verification, enabling scalable data generation while maintaining high annotation quality. We further design a unified evaluation protocol with complementary metrics tailored for null-set rejection, positive-query localization, and end-to-end GMR performance. Finally, we establish strong baselines across two modeling paradigms: a lightweight plug-and-play GMR adapter for discriminative VMR models, and a GMR-tailored GRPO reward for fine-tuning multimodal large language models (MLLMs). Extensive experiments show consistent gains across all metrics and expose key limitations of current methods, positioning GMR as a more realistic and challenging benchmark for video-language understanding.

📄 PDF Abstract BibTeX arXiv:2605.02623

Code (0)

등록된 구현이 없습니다.

Tasks

Moment Retrieval

Similar Papers 제목 키워드 기반

Finding Moments in Video Collections Using Natural Language

2019-07-30 · Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem 외

We introduce the task of retrieving relevant video moments from a large corpus of untrimmed, unsegmented videos given a natural language query. Our task poses unique challenges as a system must efficiently identify both …

Moment RetrievalRe-RankingRetrievalTemporal Localization+1

Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language

2019-12-08 · Songyang Zhang, Houwen Peng, Jianlong Fu, Jiebo Luo

We address the problem of retrieving a specific moment from an untrimmed video by a query sentence. This is a challenging problem because a target moment may take place in relations to other temporal moments in the untri…

Sentence

Joint Searching and Grounding: Multi-Granularity Video Content Retrieval

2023-10-23 · Conference 2023 10 · Zhiguo Chen, Xun Jiang, Xing Xu, Zuo Cao 외

Text-based video retrieval is a well-studied task aimed at retrieving relevant videos from a large collection in response to a given text query. Most existing TVR works assume that videos are already trimmed and fully re…

Contrastive LearningRetrievalVideo Retrieval

SemanticMoments: Training-Free Motion Similarity via Third Moment Features

2026-02-09 · Saar Huberman, Kfir Goldberg, Or Patashnik, Sagie Benaim 외 arxiv

Retrieving videos based on semantic motion is a fundamental, yet unsolved, problem. Existing video representation approaches overly rely on static appearance and scene context rather than motion dynamics, a bias inherite…

Multi-Scale 2D Temporal Adjacent Networks for Moment Localization with Natural Language

2020-12-04 · Songyang Zhang, Houwen Peng, Jianlong Fu, Yijuan Lu 외

We address the problem of retrieving a specific moment from an untrimmed video by natural language. It is a challenging problem because a target moment may take place in the context of other temporal moments in the untri…