paper-with-me

홈 › Papers

Tracking Objects and Activities with Attention for Temporal Sentence Grounding

2023-02-21 · Zeyu Xiong, Daizong Liu, Pan Zhou, Jiahao Zhu

Temporal sentence grounding (TSG) aims to localize the temporal segment which is semantically aligned with a natural language query in an untrimmed video.Most existing methods extract frame-grained features or object-grained features by 3D ConvNet or detection network under a conventional TSG framework, failing to capture the subtle differences between frames or to model the spatio-temporal behavior of core persons/objects. In this paper, we introduce a new perspective to address the TSG task by tracking pivotal objects and activities to learn more fine-grained spatio-temporal behaviors. Specifically, we propose a novel Temporal Sentence Tracking Network (TSTNet), which contains (A) a Cross-modal Targets Generator to generate multi-modal templates and search space, filtering objects and activities, and (B) a Temporal Sentence Tracker to track multi-modal targets for modeling the targets' behavior and to predict query-related segment. Extensive experiments and comparisons with state-of-the-arts are conducted on challenging benchmarks: Charades-STA and TACoS. And our TSTNet achieves the leading performance with a considerable real-time speed.

📄 PDF Abstract BibTeX arXiv:2302.10813

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceTemporal Sentence Grounding

Similar Papers 제목 키워드 기반

Video Synopsis Generation Using Spatio-Temporal Groups

2017-09-15 · A. Ahmed, D. P. Dogra, S. Kar, R. Patnaik 외

Millions of surveillance cameras operate at 24x7 generating huge amount of visual data for processing. However, retrieval of important activities from such a large data can be time consuming. Thus, researchers are workin…

ClusteringRetrievalVideo Synopsis

OVC-Net: Object-Oriented Video Captioning with Temporal Graph and Detail Enhancement

2020-03-08 · Fangyi Zhu, Jenq-Neng Hwang, Zhanyu Ma, Guang Chen 외

Traditional video captioning requests a holistic description of the video, yet the detailed descriptions of the specific objects may not be available. Without associating the moving trajectories, these image-based data-d…

ObjectSentenceVideo Captioning

DORi: Discovering Object Relationship for Moment Localization of a Natural-Language Query in Video

2020-10-13 · Cristian Rodriguez-Opazo, Edison Marrese-Taylor, Basura Fernando, Hongdong Li 외

This paper studies the task of temporal moment localization in a long untrimmed video using natural language query. Given a query sentence, the goal is to determine the start and end of the relevant segment within the vi…

Sentence

Cross-Sentence Temporal and Semantic Relations in Video Activity Localisation

2021-07-23 · ICCV 2021 10 · Jiabo Huang, Yang Liu, Shaogang Gong, Hailin Jin

Video activity localisation has recently attained increasing attention due to its practical values in automatically localising the most salient visual segments corresponding to their language descriptions (sentences) fro…

Sentence

Looking Beyond Two Frames: End-to-End Multi-Object Tracking Using Spatial and Temporal Transformers

2021-03-27 · Tianyu Zhu, Markus Hiller, Mahsa Ehsanpour, Rongkai Ma 외

Tracking a time-varying indefinite number of objects in a video sequence over time remains a challenge despite recent advances in the field. Most existing approaches are not able to properly handle multi-object tracking …

Multi-Object TrackingObjectObject TrackingOnline Multi-Object Tracking