paper-with-me

홈 › Papers

PatchTrack: Multiple Object Tracking Using Frame Patches

2022-01-01 · Xiaotong Chen, Seyed Mehdi Iranmanesh, Kuo-Chin Lien

Object motion and object appearance are commonly used information in multiple object tracking (MOT) applications, either for associating detections across frames in tracking-by-detection methods or direct track predictions for joint-detection-and-tracking methods. However, not only are these two types of information often considered separately, but also they do not help optimize the usage of visual information from the current frame of interest directly. In this paper, we present PatchTrack, a Transformer-based joint-detection-and-tracking system that predicts tracks using patches of the current frame of interest. We use the Kalman filter to predict the locations of existing tracks in the current frame from the previous frame. Patches cropped from the predicted bounding boxes are sent to the Transformer decoder to infer new tracks. By utilizing both object motion and object appearance information encoded in patches, the proposed method pays more attention to where new tracks are more likely to occur. We show the effectiveness of PatchTrack on recent MOT benchmarks, including MOT16 (MOTA 73.71%, IDF1 65.77%) and MOT17 (MOTA 73.59%, IDF1 65.23%). The results are published on https://motchallenge.net/method/MOT=4725&chl=10.

📄 PDF Abstract BibTeX arXiv:2201.00080

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderMultiple Object TrackingObjectObject Tracking

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Non-Rigid Object Tracking via Deformable Patches Using Shape-Preserved KCF and Level Sets

2017-10-01 · ICCV 2017 10 · Xin Sun, Ngai-Man Cheung, Hongxun Yao, Yiluan Guo

Part-based trackers are effective in exploiting local details of the target object for robust tracking. In contrast to most existing part-based methods that divide all kinds of target objects into a number of fixed recta…

Object Tracking

Learning Hierarchical Features for Visual Object Tracking with Recursive Neural Networks

2018-01-06 · Li Wang, Ting Liu, Bing Wang, Xulei Yang 외

Recently, deep learning has achieved very promising results in visual object tracking. Deep neural networks in existing tracking methods require a lot of training data to learn a large number of parameters. However, trai…

ObjectObject TrackingVisual Object Tracking

Reliable Patch Trackers: Robust Visual Tracking by Exploiting Reliable Patches

2015-06-01 · CVPR 2015 6 · Yang Li, Jianke Zhu, Steven C. H. Hoi

Most modern trackers typically employ a bounding box given in the first frame to track visual objects, where their tracking results are often sensitive to the initialization. In this paper, we propose a new tracking meth…

ClusteringVisual Tracking

SOWP: Spatially Ordered and Weighted Patch Descriptor for Visual Tracking

2015-12-01 · ICCV 2015 12 · Han-Ul Kim, Dae-Youn Lee, Jae-Young Sim, Chang-Su Kim

A simple yet effective object descriptor for visual tracking is proposed in this paper. We first decompose the bounding box of a target object into multiple patches, which are described by color and gradient histograms. …

ObjectVisual Tracking

Part-based Tracking by Sampling

2018-05-22 · George De Ath, Richard M. Everson

We propose a novel part-based method for tracking an arbitrary object in challenging video sequences. The colour distribution of tracked image patches on the target object are represented by pairs of RGB samples and coun…

Object