paper-with-me

홈 › Papers

Video Moment Retrieval from Text Queries via Single Frame Annotation

2022-04-20 · Ran Cui, Tianwen Qian, Pai Peng, Elena Daskalaki, Jingjing Chen, Xiaowei Guo, Huyang Sun, Yu-Gang Jiang

Video moment retrieval aims at finding the start and end timestamps of a moment (part of a video) described by a given natural language query. Fully supervised methods need complete temporal boundary annotations to achieve promising results, which is costly since the annotator needs to watch the whole moment. Weakly supervised methods only rely on the paired video and query, but the performance is relatively poor. In this paper, we look closer into the annotation process and propose a new paradigm called "glance annotation". This paradigm requires the timestamp of only one single random frame, which we refer to as a "glance", within the temporal boundary of the fully supervised counterpart. We argue this is beneficial because comparing to weak supervision, trivial cost is added yet more potential in performance is provided. Under the glance annotation setting, we propose a method named as Video moment retrieval via Glance Annotation (ViGA) based on contrastive learning. ViGA cuts the input video into clips and contrasts between clips and queries, in which glance guided Gaussian distributed weights are assigned to all clips. Our extensive experiments indicate that ViGA achieves better results than the state-of-the-art weakly supervised methods by a large margin, even comparable to fully supervised methods in some cases.

📄 PDF Abstract BibTeX arXiv:2204.09409

Code (1)

r-cui/ViGA 공식 구현 pytorch

Tasks

Contrastive LearningMoment RetrievalRetrieval

Similar Papers 제목 키워드 기반

MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval

2025-02-18 · Huaying Yuan, Jian Ni, Zheng Liu, Yueze Wang 외

Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or the…

Action RecognitionMoment RetrievalObject LocalizationRAG+2

Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval Using Language

2026-05-28 · Xiang Fang, Wanlong Fang, Daizong Liu, Xiaoye Qu 외 arxiv

Video Moment Retrieval (VMR) targets to retrieve the specific moment corresponding to a sentence query from an untrimmed video. Although recent works have made remarkable progress in this task, they implicitly are rooted…

Activity DetectionMoment Retrieval

VeRVE: Versatile Retrieval for Videos via Unified Embeddings

2026-01-17 · Shaunak Halbe, Bhagyashree Puranik, Jayakrishnan Unnikrishnan, Kushan Thakkar 외 arxiv

Modern video retrieval systems are expected to handle diverse tasks ranging from corpus-level retrieval, fine-grained moment localization to flexible multimodal querying. Specialized architectures achieve strong retrieva…

Zero-shot Moment RetrievalZero-Shot Video Retrieval

Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval

2025-02-12 · Kevin Flanagan, Dima Damen, Michael Wray

Video Moment Retrieval is a common task to evaluate the performance of visual-language models - it involves localising start and end times of moments in videos from query sentences. The current task formulation assumes t…

AvgMoment RetrievalRetrieval

Towards Efficient Partially Relevant Video Retrieval with Active Moment Discovering

2025-04-15 · Peipei Song, Long Zhang, Long Lan, Weidong Chen 외

Partially relevant video retrieval (PRVR) is a practical yet challenging task in text-to-video retrieval, where videos are untrimmed and contain much background content. The pursuit here is of both effective and efficien…

Partially Relevant Video RetrievalRetrievalText to Video RetrievalVideo Retrieval