Learning to Track Objects from Unlabeled Videos
In this paper, we propose to learn an Unsupervised Single Object Tracker (USOT) from scratch. We identify that three major challenges, i.e., moving object discovery, rich temporal variation exploitation, and online update, are the central causes of the performance bottleneck of existing unsupervised trackers. To narrow the gap between unsupervised trackers and supervised counterparts, we propose an effective unsupervised learning approach composed of three stages. First, we sample sequentially moving objects with unsupervised optical flow and dynamic programming, instead of random cropping. Second, we train a naive Siamese tracker from scratch using single-frame pairs. Third, we continue training the tracker with a novel cycle memory learning scheme, which is conducted in longer temporal spans and also enables our tracker to update online. Extensive experiments show that the proposed USOT learned from unlabeled videos performs well over the state-of-the-art unsupervised trackers by large margins, and on par with recent supervised deep trackers. Code is available at https://github.com/VISION-SJTU/USOT.
Code (1)
Tasks
Object DiscoveryOptical Flow EstimationSimilar Papers 제목 키워드 기반
Semi-TCL: Semi-Supervised Track Contrastive Representation Learning
Online tracking of multiple objects in videos requires strong capacity of modeling and matching object appearances. Previous methods for learning appearance embedding mostly rely on instance-level matching without consid…
Multiple Object TrackingObjectObject TrackingRepresentation LearningLearning to Separate Object Sounds by Watching Unlabeled Video
Perceiving a scene most fully requires all the senses. Yet modeling how objects look and sound is challenging: most natural scenes and events contain multiple objects, and the audio track mixes all the sound sources toge…
Audio DenoisingAudio Source SeparationDenoisingMulti-Label LearningTwo Video Data Sets for Tracking and Retrieval of Out of Distribution Objects
In this work we present two video test data sets for the novel computer vision (CV) task of out of distribution tracking (OOD tracking). Here, OOD objects are understood as objects with a semantic class outside the seman…
Image SegmentationRetrievalSemantic SegmentationSelf-supervised Moving Vehicle Tracking with Stereo Sound
Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabele…
Object Localizationvehicle detectionVisual LocalizationMulti-Object Tracking with Hallucinated and Unlabeled Videos
In this paper, we explore learning end-to-end deep neural trackers without tracking annotations. This is important as large-scale training data is essential for training deep neural trackers while tracking annotations ar…
Multi-Object TrackingObjectObject Tracking