Joint Learning of Siamese CNNs and Temporally Constrained Metrics for Tracklet Association
In this paper, we study the challenging problem of multi-object tracking in a complex scene captured by a single camera. Different from the existing tracklet association-based tracking methods, we propose a novel and efficient way to obtain discriminative appearance-based tracklet affinity models. Our proposed method jointly learns the convolutional neural networks (CNNs) and temporally constrained metrics. In our method, a Siamese convolutional neural network (CNN) is first pre-trained on the auxiliary data. Then the Siamese CNN and temporally constrained metrics are jointly learned online to construct the appearance-based tracklet affinity models. The proposed method can jointly learn the hierarchical deep features and temporally constrained segment-wise metrics under a unified framework. For reliable association between tracklets, a novel loss function incorporating temporally constrained multi-task learning mechanism is proposed. By employing the proposed method, tracklet association can be accomplished even in challenging situations. Moreover, a new dataset with 40 fully annotated sequences is created to facilitate the tracking evaluation. Experimental results on five public datasets and the new large-scale dataset show that our method outperforms several state-of-the-art approaches in multi-object tracking.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-Object TrackingMulti-Task LearningObject TrackingSimilar Papers 제목 키워드 기반
Extending Multi-Object Tracking systems to better exploit appearance and 3D information
Tracking multiple objects in real time is essential for a variety of real-world applications, with self-driving industry being at the foremost. This work involves exploiting temporally varying appearance and motion infor…
Multi-Object TrackingObjectObject TrackingReal-Time Multi-Object TrackingDeep Learning Architectures for Code-Modulated Visual Evoked Potentials Detection
Non-invasive Brain-Computer Interfaces (BCIs) based on Code-Modulated Visual Evoked Potentials (C-VEPs) require highly robust decoding methods to address temporal variability and session-dependent noise in EEG signals. T…
Data AugmentationDeep Siamese Networks with Bayesian non-Parametrics for Video Object Tracking
We present a novel algorithm utilizing a deep Siamese neural network as a general object similarity function in combination with a Bayesian optimization (BO) framework to encode spatio-temporal information for efficient …
Bayesian OptimizationObjectObject TrackingVideo Object TrackingHierarchical Attention Diffusion Networks with Object Priors for Video Change Detection
We present a unified change detection pipeline that combines instance level masking, multi\-scale attention within a denoising diffusion model, and per pixel semantic classification, all refined via SSIM to match human p…
Change DetectionDenoisingSSIMSiamese Network for RGB-D Salient Object Detection and Beyond
Existing RGB-D salient object detection (SOD) models usually treat RGB and depth as independent information and design separate networks for feature extraction from each. Such schemes can easily be constrained by a limit…
object-detectionObject DetectionRGB-D Salient Object DetectionRGB Salient Object Detection+2