Tracking Objects and Activities with Attention for Temporal Sentence Grounding
Temporal sentence grounding (TSG) aims to localize the temporal segment which is semantically aligned with a natural language query in an untrimmed video.Most existing methods extract frame-grained features or object-grained features by 3D ConvNet or detection network under a conventional TSG framework, failing to capture the subtle differences between frames or to model the spatio-temporal behavior of core persons/objects. In this paper, we introduce a new perspective to address the TSG task by tracking pivotal objects and activities to learn more fine-grained spatio-temporal behaviors. Specifically, we propose a novel Temporal Sentence Tracking Network (TSTNet), which contains (A) a Cross-modal Targets Generator to generate multi-modal templates and search space, filtering objects and activities, and (B) a Temporal Sentence Tracker to track multi-modal targets for modeling the targets' behavior and to predict query-related segment. Extensive experiments and comparisons with state-of-the-arts are conducted on challenging benchmarks: Charades-STA and TACoS. And our TSTNet achieves the leading performance with a considerable real-time speed.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceTemporal Sentence GroundingSimilar Papers 제목 키워드 기반
Video Synopsis Generation Using Spatio-Temporal Groups
Millions of surveillance cameras operate at 24x7 generating huge amount of visual data for processing. However, retrieval of important activities from such a large data can be time consuming. Thus, researchers are workin…
ClusteringRetrievalVideo SynopsisOVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement
Traditional video captioning requests a holistic description of the video, yet the detailed descriptions of the specific objects may not be available. Without associating the moving trajectories, these image-based data-d…
ObjectSentenceVideo CaptioningDORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video
This paper studies the task of temporal moment localization in a long untrimmed video using natural language query. Given a query sentence, the goal is to determine the start and end of the relevant segment within the vi…
SentenceCross-Sentence Temporal and Semantic Relations in Video Activity Localisation
Video activity localisation has recently attained increasing attention due to its practical values in automatically localising the most salient visual segments corresponding to their language descriptions (sentences) fro…
SentenceLooking Beyond Two Frames: End-to-End Multi-Object Tracking Using Spatial and Temporal Transformers
Tracking a time-varying indefinite number of objects in a video sequence over time remains a challenge despite recent advances in the field. Most existing approaches are not able to properly handle multi-object tracking …
Multi-Object TrackingObjectObject TrackingOnline Multi-Object Tracking