paper-with-me

홈 › Papers

VideoTrack: Learning To Track Objects via Video Transformer

2023-01-01 · CVPR 2023 1 · Fei Xie, Lei Chu, Jiahao Li, Yan Lu, Chao Ma

Existing Siamese tracking methods, which are built on pair-wise matching between two single frames, heavily rely on additional sophisticated mechanism to exploit temporal information among successive video frames, hindering them from high efficiency and industrial deployments. In this work, we resort to sequence-level target matching that can encode temporal contexts into the spatial features through a neat feedforward video model. Specifically, we adapt the standard video transformer architecture to visual tracking by enabling spatiotemporal feature learning directly from frame-level patch sequences. To better adapt to the tracking task, we carefully blend the spatiotemporal information in the video clips through sequential multi-branch triplet blocks, which formulates a video transformer backbone. Our experimental study compares different model variants, such as tokenization strategies, hierarchical structures, and video attention schemes. Then, we propose a disentangled dual-template mechanism that decouples static and dynamic appearance changes over time, and reduces the temporal redundancy in video frames. Extensive experiments show that our method, named as VideoTrack, achieves state-of-the-art results while running in real-time.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

TripletVisual Tracking

Similar Papers 제목 키워드 기반

TransMOT: Spatial-Temporal Graph Transformer for Multiple Object Tracking

2021-04-01 · Peng Chu, Jiang Wang, Quanzeng You, Haibin Ling 외

Tracking multiple objects in videos relies on modeling the spatial-temporal interactions of the objects. In this paper, we propose a solution named TransMOT, which leverages powerful graph transformers to efficiently mod…

DecoderMulti-Object TrackingMultiple Object TrackingObject+2

Temporal-Enhanced Multimodal Transformer for Referring Multi-Object Tracking and Segmentation

2024-10-17 · Changcheng Xiao, Qiong Cao, Yujie Zhong, Xiang Zhang 외

Referring multi-object tracking (RMOT) is an emerging cross-modal task that aims to locate an arbitrary number of target objects and maintain their identities referred by a language expression in a video. This intricate …

Multi-Object TrackingMulti-Object Tracking and SegmentationObject TrackingReferring Multi-Object Tracking+3

Whareformer: Learning to Track What is Where in Long Egocentric Videos

2026-07-09 · Jacob Chalk, Saptarshi Sinha, Dima Damen, Yannis Kalantidis 외 arxiv

The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout …

Track Targets by Dense Spatio-Temporal Position Encoding

2022-10-17 · Jinkun Cao, Hao Wu, Kris Kitani

In this work, we propose a novel paradigm to encode the position of targets for target tracking in videos using transformers. The proposed paradigm, Dense Spatio-Temporal (DST) position encoding, encodes spatio-temporal …

Multi-Object TrackingMulti-Object Tracking and SegmentationObjectObject Tracking+1

EchoTrack: Auditory Referring Multi-Object Tracking for Autonomous Driving

2024-02-28 · Jiacheng Lin, Jiajun Chen, Kunyu Peng, Xuan He 외

This paper introduces the task of Auditory Referring Multi-Object Tracking (AR-MOT), which dynamically tracks specific objects in a video sequence based on audio expressions and appears as a challenging problem in autono…

Autonomous DrivingMulti-Object TrackingObjectObject Tracking+1