paper-with-me

홈 › Papers

MAN: Moment Alignment Network for Natural Language Moment Retrieval via Iterative Graph Adjustment

2018-11-30 · CVPR 2019 6 · Da Zhang, Xiyang Dai, Xin Wang, Yuan-Fang Wang, Larry S. Davis

This research strives for natural language moment retrieval in long, untrimmed video streams. The problem is not trivial especially when a video contains multiple moments of interests and the language describes complex temporal dependencies, which often happens in real scenarios. We identify two crucial challenges: semantic misalignment and structural misalignment. However, existing approaches treat different moments separately and do not explicitly model complex moment-wise temporal relations. In this paper, we present Moment Alignment Network (MAN), a novel framework that unifies the candidate moment encoding and temporal structural reasoning in a single-shot feed-forward network. MAN naturally assigns candidate moment representations aligned with language semantics over different temporal locations and scales. Most importantly, we propose to explicitly model moment-wise temporal relations as a structured graph and devise an iterative graph adjustment network to jointly learn the best structure in an end-to-end manner. We evaluate the proposed approach on two challenging public benchmarks DiDeMo and Charades-STA, where our MAN significantly outperforms the state-of-the-art by a large margin.

📄 PDF Abstract BibTeX arXiv:1812.00087

Code (0)

등록된 구현이 없습니다.

Tasks

Moment RetrievalNatural Language Moment RetrievalRetrieval

Similar Papers 제목 키워드 기반

Finding Moments in Video Collections Using Natural Language

2019-07-30 · Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem 외

We introduce the task of retrieving relevant video moments from a large corpus of untrimmed, unsegmented videos given a natural language query. Our task poses unique challenges as a system must efficiently identify both …

Moment RetrievalRe-RankingRetrievalTemporal Localization+1

Background-aware Moment Detection for Video Moment Retrieval

2023-06-05 · Minjoon Jung, Youwon Jang, SeongHo Choi, Joochan Kim 외

Video moment retrieval (VMR) identifies a specific moment in an untrimmed video for a given natural language query. This task is prone to suffer the weak alignment problem innate in video datasets. Due to the ambiguity, …

Moment RetrievalNatural Language Moment RetrievalRetrieval

Disentangle and denoise: Tackling context misalignment for video moment retrieval

2024-08-14 · Kaijing Ma, Han Fang, Xianghao Zang, Chao Ban 외

Video Moment Retrieval, which aims to locate in-context video moments according to a natural language query, is an essential task for cross-modal grounding. Existing methods focus on enhancing the cross-modal interaction…

DenoisingDisentanglementMoment RetrievalRetrieval+1

Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment Retrieval

2025-11-30 · Haojian Huang, Kaijing Ma, Jin Chen, Haodong Chen 외 arxiv

In the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine…

Natural Language QueriesMoment Retrieval

GenSpan: Generation-Calibrated Motion Span Priors for Multi-Verb Video Corpus Moment Retrieval

2026-03-23 · Yunzhuo Sun, Xinyue Liu, Yanyang Li, Nanding Wu 외 arxiv

Video Corpus Moment Retrieval (VCMR) aims to retrieve both the correct video and its temporal segment corresponding to a natural-language query, a task that is especially challenging for multi-verb queries where temporal…

Moment Retrieval