Learning a Deep Compact Image Representation for Visual Tracking
In this paper, we study the challenging problem of tracking the trajectory of a moving object in a video with possibly very complex background. In contrast to most existing trackers which only learn the appearance of the tracked object online, we take a different approach, inspired by recent advances in deep learning architectures, by putting more emphasis on the (unsupervised) feature learning problem. Specifically, by using auxiliary natural images, we train a stacked denoising autoencoder offline to learn generic image features that are more robust against variations. This is then followed by knowledge transfer from offline training to the online tracking process. Online tracking involves a classification neural network which is constructed from the encoder part of the trained autoencoder as a feature extractor and an additional classification layer. Both the feature extractor and the classifier can be further tuned to adapt to appearance changes of the moving object. Comparison with the state-of-the-art trackers on some challenging benchmark video sequences shows that our deep learning tracker is very efficient as well as more accurate.
Code (0)
등록된 구현이 없습니다.
Tasks
DenoisingGeneral ClassificationObjectTransfer LearningVisual TrackingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Compact Binary Codes for Visual Tracking
A key problem in visual tracking is to represent the appearance of an object in a way that is robust to visual changes. To attain this robustness, increasingly complex models are used to capture appearance variations. Ho…
Visual TrackingLearning Compact Target-Oriented Feature Representations for Visual Tracking
Many state-of-the-art trackers usually resort to the pretrained convolutional neural network (CNN) model for correlation filtering, in which deep features could usually be redundant, noisy and less discriminative for som…
Visual TrackingOnline Hybrid Lightweight Representations Learning: Its Application to Visual Tracking
This paper presents a novel hybrid representation learning framework for streaming data, where an image frame in a video is modeled by an ensemble of two distinct deep neural networks; one is a low-bit quantized network …
Representation LearningVisual TrackingCompact Transformer Tracker with Correlative Masked Modeling
Transformer framework has been showing superior performances in visual object tracking for its great strength in information aggregation across the template and search image with the well-known attention mechanism. Most …
DecoderObject TrackingVisual Object TrackingMotionGS : Compact Gaussian Splatting SLAM by Motion Filter
With their high-fidelity scene representation capability, the attention of SLAM field is deeply attracted by the Neural Radiation Field (NeRF) and 3D Gaussian Splatting (3DGS). Recently, there has been a surge in NeRF-ba…
3DGSNeRFPose Estimation