paper-with-me

Papers

Parallel Attention Network with Sequence Matching for Video Grounding

2021-05-18 · Findings (ACL) 2021 8 · Hao Zhang, Aixin Sun, Wei Jing, Liangli Zhen, Joey Tianyi Zhou, Rick Siow Mong Goh

Given a video, video grounding aims to retrieve a temporal moment that semantically corresponds to a language query. In this work, we propose a Parallel Attention Network with Sequence matching (SeqPAN) to address the challenges in this task: multi-modal representation learning, and target moment boundary prediction. We design a self-guided parallel attention module to effectively capture self-modal contexts and cross-modal attentive information between video and text. Inspired by sequence labeling tasks in natural language processing, we split the ground truth moment into begin, inside, and end regions. We then propose a sequence matching strategy to guide start/end boundary predictions using region labels. Experimental results on three datasets show that SeqPAN is superior to state-of-the-art methods. Furthermore, the effectiveness of the self-guided parallel attention module and the sequence matching module is verified.

📄 PDF Abstract BibTeX arXiv:2105.08481

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningVideo Grounding

Similar Papers 제목 키워드 기반

End-to-End Dense Video Grounding via Parallel Regression

2021-09-23 · Fengyuan Shi, Weilin Huang, LiMin Wang

Video grounding aims to localize the corresponding video moment in an untrimmed video given a language query. Existing methods often address this task in an indirect way, by casting it as a proposal-and-match or fusion-a…

regressionSentenceVideo Grounding

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

2026-08-28 · Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan 외 arxiv

Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense…

Spatio-Temporal Video GroundingVideo Object Tracking

VLG-Net: Video-Language Graph Matching Network for Video Grounding

2020-11-19 · Mattia Soldan, Mengmeng Xu, Sisi Qu, Jesper Tegner 외

Grounding language queries in videos aims at identifying the time interval (or moment) semantically relevant to a language query. The solution to this challenging task demands understanding videos' and queries' semantic …

Graph MatchingMoment RetrievalNatural Language Moment RetrievalTemporal Localization+1

Co-Grounding Networks with Semantic Attention for Referring Expression Comprehension in Videos

2021-03-23 · CVPR 2021 1 · Sijie Song, Xudong Lin, Jiaying Liu, Zongming Guo 외

In this paper, we address the problem of referring expression comprehension in videos, which is challenging due to complex expression and scene dynamics. Unlike previous methods which solve the problem in multiple stages…

Referring ExpressionReferring Expression ComprehensionVideo Grounding

MRTNet: Multi-Resolution Temporal Network for Video Sentence Grounding

2022-12-26 · Wei Ji, Long Chen, Yinwei Wei, Yiming Wu 외

Given an untrimmed video and natural language query, video sentence grounding aims to localize the target temporal moment in the video. Existing methods mainly tackle this task by matching and aligning semantics of the d…

DecoderDescriptiveSentence