VideoTrack: Learning To Track Objects via Video Transformer
Existing Siamese tracking methods, which are built on pair-wise matching between two single frames, heavily rely on additional sophisticated mechanism to exploit temporal information among successive video frames, hindering them from high efficiency and industrial deployments. In this work, we resort to sequence-level target matching that can encode temporal contexts into the spatial features through a neat feedforward video model. Specifically, we adapt the standard video transformer architecture to visual tracking by enabling spatiotemporal feature learning directly from frame-level patch sequences. To better adapt to the tracking task, we carefully blend the spatiotemporal information in the video clips through sequential multi-branch triplet blocks, which formulates a video transformer backbone. Our experimental study compares different model variants, such as tokenization strategies, hierarchical structures, and video attention schemes. Then, we propose a disentangled dual-template mechanism that decouples static and dynamic appearance changes over time, and reduces the temporal redundancy in video frames. Extensive experiments show that our method, named as VideoTrack, achieves state-of-the-art results while running in real-time.
Code (0)
등록된 구현이 없습니다.
Tasks
TripletVisual TrackingSimilar Papers 제목 키워드 기반
TransMOT: Spatial-Temporal Graph Transformer for Multiple Object Tracking
Tracking multiple objects in videos relies on modeling the spatial-temporal interactions of the objects. In this paper, we propose a solution named TransMOT, which leverages powerful graph transformers to efficiently mod…
DecoderMulti-Object TrackingMultiple Object TrackingObject+2Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and Segmentation
Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a language expression in a video. This intricate …
Multi-Object TrackingMulti-Object Tracking and SegmentationObject TrackingReferring Multi-Object Tracking+3Whareformer: Learning to Track What is Where in Long Egocentric Videos
The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout …
Track Targets by Dense Spatio-Temporal Position Encoding
In this work, we propose a novel paradigm to encode the position of targets for target tracking in videos using transformers. The proposed paradigm, Dense Spatio-Temporal (DST) position encoding, encodes spatio-temporal …
Multi-Object TrackingMulti-Object Tracking and SegmentationObjectObject Tracking+1EchoTrack: Auditory Referring Multi-Object Tracking for Autonomous Driving
This paper introduces the task of Auditory Referring Multi-Object Tracking (AR-MOT), which dynamically tracks specific objects in a video sequence based on audio expressions and appears as a challenging problem in autono…
Autonomous DrivingMulti-Object TrackingObjectObject Tracking+1