Parallel Attention Network with Sequence Matching for Video Grounding
Given a video, video grounding aims to retrieve a temporal moment that semantically corresponds to a language query. In this work, we propose a Parallel Attention Network with Sequence matching (SeqPAN) to address the challenges in this task: multi-modal representation learning, and target moment boundary prediction. We design a self-guided parallel attention module to effectively capture self-modal contexts and cross-modal attentive information between video and text. Inspired by sequence labeling tasks in natural language processing, we split the ground truth moment into begin, inside, and end regions. We then propose a sequence matching strategy to guide start/end boundary predictions using region labels. Experimental results on three datasets show that SeqPAN is superior to state-of-the-art methods. Furthermore, the effectiveness of the self-guided parallel attention module and the sequence matching module is verified.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation LearningVideo GroundingSimilar Papers 제목 키워드 기반
End-to-End Dense Video Grounding via Parallel Regression
Video grounding aims to localize the corresponding video moment in an untrimmed video given a language query. Existing methods often address this task in an indirect way, by casting it as a proposal-and-match or fusion-a…
regressionSentenceVideo GroundingLocate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense…
Spatio-Temporal Video GroundingVideo Object TrackingVLG-Net: Video-Language Graph Matching Network for Video Grounding
Grounding language queries in videos aims at identifying the time interval (or moment) semantically relevant to a language query. The solution to this challenging task demands understanding videos' and queries' semantic …
Graph MatchingMoment RetrievalNatural Language Moment RetrievalTemporal Localization+1Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos
In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages…
Referring ExpressionReferring Expression ComprehensionVideo GroundingMRTNet: Multi-Resolution Temporal Network for Video Sentence Grounding
Given an untrimmed video and natural language query, video sentence grounding aims to localize the target temporal moment in the video. Existing methods mainly tackle this task by matching and aligning semantics of the d…
DecoderDescriptiveSentence