Span-based Localizing Network for Natural Language Video Localization
Given an untrimmed video and a text query, natural language video localization (NLVL) is to locate a matching span from the video that semantically corresponds to the query. Existing solutions formulate NLVL either as a ranking task and apply multimodal matching architecture, or as a regression task to directly regress the target video span. In this work, we address NLVL task with a span-based QA approach by treating the input video as text passage. We propose a video span localizing network (VSLNet), on top of the standard span-based QA framework, to address NLVL. The proposed VSLNet tackles the differences between NLVL and span-based QA through a simple yet effective query-guided highlighting (QGH) strategy. The QGH guides VSLNet to search for matching video span within a highlighted region. Through extensive experiments on three benchmark datasets, we show that the proposed VSLNet outperforms the state-of-the-art methods; and adopting span-based QA framework is a promising direction to solve NLVL.
Code (1)
Tasks
Temporal Sentence GroundingSimilar Papers 제목 키워드 기반
Natural Language Video Localization: A Revisit in Span-based Question Answering Framework
Natural Language Video Localization (NLVL) aims to locate a target moment from an untrimmed video that semantically corresponds to a text query. Existing approaches mainly solve the NLVL problem from the perspective of c…
Question AnsweringLocalizing Events in Videos with Multimodal Queries
Localizing events in videos based on semantic queries is a pivotal task in video understanding, with the growing significance of user-oriented applications like video search. Yet, current research predominantly relies on…
Natural Language QueriesVideo UnderstandingLocalizing Moments in Video with Temporal Language
Localizing moments in a longer video via natural language queries is a new, challenging task at the intersection of language and video understanding. Though moment localization with natural language is similar to other l…
Natural Language QueriesRetrievalVideo UnderstandingUAL-Bench: The First Comprehensive Unusual Activity Localization Benchmark
Localizing unusual activities, such as human errors or surveillance incidents, in videos holds practical significance. However, current video understanding models struggle with localizing these unusual events likely beca…
Unusual Activity LocalizationVideo UnderstandingPrompting Large Language Models to Reformulate Queries for Moment Localization
The task of moment localization is to localize a temporal moment in an untrimmed video for a given natural language query. Since untrimmed video contains highly redundant contents, the quality of the query is crucial for…
Moment QueriesNatural Language Queries