paper-with-me

홈 › Papers

MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval

2025-02-18 · Huaying Yuan, Jian Ni, Zheng Liu, Yueze Wang, Junjie Zhou, Zhengyang Liang, Bo Zhao, Zhao Cao, Zhicheng Dou, Ji-Rong Wen

Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on the end-to-end LVU performance, making them inappropriate for evaluating whether key moments can be accurately accessed. To address this challenge, we propose MomentSeeker, a novel benchmark for long-video moment retrieval (LMVR), distinguished by the following features. First, it is created based on long and diverse videos, averaging over 1200 seconds in duration and collected from various domains, e.g., movie, anomaly, egocentric, and sports. Second, it covers a variety of real-world scenarios in three levels: global-level, event-level, object-level, covering common tasks like action recognition, object localization, and causal reasoning, etc. Third, it incorporates rich forms of queries, including text-only queries, image-conditioned queries, and video-conditioned queries. On top of MomentSeeker, we conduct comprehensive experiments for both generation-based approaches (directly using MLLMs) and retrieval-based approaches (leveraging video retrievers). Our results reveal the significant challenges in long-video moment retrieval in terms of accuracy and efficiency, despite improvements from the latest long-video MLLMs and task-specific fine-tuning. We have publicly released MomentSeeker(https://yhy-2000.github.io/MomentSeeker/) to facilitate future research in this area.

📄 PDF Abstract BibTeX arXiv:2502.12558

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionMoment RetrievalObject LocalizationRAGRetrievalVideo Understanding

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Towards Event-oriented Long Video Understanding

2024-06-20 · Yifan Du, Kun Zhou, Yuqi Huo, YiFan Li 외

With the rapid development of video Multimodal Large Language Models (MLLMs), numerous benchmarks have been proposed to assess their video understanding capability. However, due to the lack of rich events in the videos, …

Video Understanding

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

2026-08-13 · Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi 외 arxiv

Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit. However, existing benchmarks r…

Speaker Verification

SOVC: Subject-Oriented Video Captioning

2023-12-20 · Chang Teng, Yunchuan Ma, Guorong Li, Yuankai Qi 외

Describing video content according to users' needs is a long-held goal. Although existing video captioning methods have made significant progress, the generated captions may not focus on the entity that users are particu…

Video Captioning

Video Enhancement with Task-Oriented Flow

2017-11-24 · Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei 외

Many video enhancement algorithms rely on optical flow to register frames in a video sequence. Precise flow estimation is however intractable; and optical flow itself is often a sub-optimal representation for particular …

DenoisingMotion EstimationOptical Flow EstimationSuper-Resolution+4

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

2026-05-08 · Yuancheng Wei, Linli Yao, Lei Li, Haojie Zhang 외 arxiv

Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust evaluation benchmarks and high-quality pr…