paper-with-me

홈 › Papers

Object-Centric Framework for Video Moment Retrieval

2025-12-20 · Zongyao Li, Yongkang Wong, Satoshi Yamazaki, Jianquan Liu, Mohan Kankanhalli arxiv

Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing moments described by object-oriented queries involving specific entities and their interactions. In particular, temporal dynamics at the object level have been largely overlooked, limiting the effectiveness of existing approaches in scenarios requiring detailed object-level reasoning. To address this limitation, we propose a novel object-centric framework for moment retrieval. Our method first extracts query-relevant objects using a scene graph parser and then generates scene graphs from video frames to represent these objects and their relationships. Based on the scene graphs, we construct object-level feature sequences that encode rich visual and semantic information. These sequences are processed by a relational tracklet transformer, which models spatio-temporal correlations among objects over time. By explicitly capturing object-level state changes, our framework enables more accurate localization of moments aligned with object-oriented queries. We evaluated our method on three benchmarks: Charades-STA, QVHighlights, and TACoS. Experimental results demonstrate that our method outperforms existing state-of-the-art methods across all benchmarks.

📄 PDF Abstract BibTeX arXiv:2512.18448

Code (0)

등록된 구현이 없습니다.

Tasks

Temporal SequencesMoment Retrieval

Similar Papers 제목 키워드 기반

MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval

2025-02-18 · Huaying Yuan, Jian Ni, Zheng Liu, Yueze Wang 외

Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or the…

Action RecognitionMoment RetrievalObject LocalizationRAG+2

Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning

2024-12-18 · Yunbin Tu, Liang Li, Li Su, Qingming Huang

Video has emerged as a favored multimedia format on the internet. To better gain video contents, a new topic HIREST is presented, including video retrieval, moment retrieval, moment segmentation, and step-captioning. The…

Moment RetrievalMulti-Task LearningRetrievalVideo Retrieval+1

Egocentric Video-Language Pretraining

2022-06-03 · Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray 외

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-sc…

Action RecognitionContrastive LearningMoment QueriesMulti-Instance Retrieval+9

Selective Query-guided Debiasing for Video Corpus Moment Retrieval

2022-10-17 · Sunjae Yoon, Ji Woo Hong, Eunseop Yoon, Dahyun Kim 외

Video moment retrieval (VMR) aims to localize target moments in untrimmed videos pertinent to a given textual query. Existing retrieval systems tend to rely on retrieval bias as a shortcut and thus, fail to sufficiently …

Moment RetrievalRetrievalVideo Corpus Moment Retrieval

EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

2024-06-26 · Baoqi Pei, Guo Chen, Jilan Xu, Yuping He 외

In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Building upon the video-language two-tower mod…

Action AnticipationAction RecognitionDomain AdaptationLong Term Action Anticipation+4