paper-with-me

홈 › Papers

Faster Video Moment Retrieval with Point-Level Supervision

2023-05-23 · Xun Jiang, Zailei Zhou, Xing Xu, Yang Yang, Guoqing Wang, Heng Tao Shen

Video Moment Retrieval (VMR) aims at retrieving the most relevant events from an untrimmed video with natural language queries. Existing VMR methods suffer from two defects: (1) massive expensive temporal annotations are required to obtain satisfying performance; (2) complicated cross-modal interaction modules are deployed, which lead to high computational cost and low efficiency for the retrieval process. To address these issues, we propose a novel method termed Cheaper and Faster Moment Retrieval (CFMR), which well balances the retrieval accuracy, efficiency, and annotation cost for VMR. Specifically, our proposed CFMR method learns from point-level supervision where each annotation is a single frame randomly located within the target moment. It is 6 times cheaper than the conventional annotations of event boundaries. Furthermore, we also design a concept-based multimodal alignment mechanism to bypass the usage of cross-modal interaction modules during the inference process, remarkably improving retrieval efficiency. The experimental results on three widely used VMR benchmarks demonstrate the proposed CFMR method establishes new state-of-the-art with point-level supervision. Moreover, it significantly accelerates the retrieval speed with more than 100 times FLOPs compared to existing approaches with point-level supervision.

📄 PDF Abstract BibTeX arXiv:2305.14017

Code (0)

등록된 구현이 없습니다.

Tasks

Moment RetrievalNatural Language QueriesRetrieval

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

2DP-2MRC: 2-Dimensional Pointer-based Machine Reading Comprehension Method for Multimodal Moment Retrieval

2024-06-10 · Jiajun He, Tomoki Toda

Moment retrieval aims to locate the most relevant moment in an untrimmed video based on a given natural language query. Existing solutions can be roughly categorized into moment-based and clip-based methods. The former o…

Boundary DetectionMachine Reading ComprehensionMoment RetrievalReading Comprehension+1

Finding Moments in Video Collections Using Natural Language

2019-07-30 · Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem 외

We introduce the task of retrieving relevant video moments from a large corpus of untrimmed, unsegmented videos given a natural language query. Our task poses unique challenges as a system must efficiently identify both …

Moment RetrievalRe-RankingRetrievalTemporal Localization+1

Language-based Audio Moment Retrieval

2024-09-24 · Hokuto Munakata, Taichi Nishimura, Shota Nakada, Tatsuya Komatsu

In this paper, we propose and design a new task called audio moment retrieval (AMR). Unlike conventional language-based audio retrieval tasks that search for short audio clips from an audio database, AMR aims to predict …

audio moment retrievalMoment RetrievalRetrieval

MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval

2025-02-18 · Huaying Yuan, Jian Ni, Zheng Liu, Yueze Wang 외

Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or the…

Action RecognitionMoment RetrievalObject LocalizationRAG+2

Hierarchical Video-Moment Retrieval and Step-Captioning

2023-03-29 · CVPR 2023 1 · Abhay Zala, Jaemin Cho, Satwik Kottur, Xilun Chen 외

There is growing interest in searching for information from large video corpora. Prior works have studied relevant tasks, such as text-based video retrieval, moment retrieval, video summarization, and video captioning in…

Information RetrievalMoment RetrievalRetrievalVideo Captioning+2