DecoderTracker: Decoder-Only Method for Multiple-Object Tracking
Decoder-only models, such as GPT, have demonstrated superior performance in many areas compared to traditional encoder-decoder structure transformer models. Over the years, end-to-end models based on the traditional transformer structure, like MOTR, have achieved remarkable performance in multi-object tracking. However, the significant computational resource consumption of these models leads to less friendly inference speeds and training times. To address these issues, this paper attempts to construct a lightweight Decoder-only model: DecoderTracker for end-to-end multi-object tracking. Specifically, drawing on some real-time detection models, we have developed an image feature extraction network which can efficiently extract features from images to replace the encoder structure. In addition to minor innovations in the network, we analyze the potential reasons for the slow training of MOTR-like models and propose an effective training strategy to mitigate the issue of prolonged training times. On the DanceTrack dataset, without any bells and whistles, DecoderTracker's tracking performance slightly surpasses that of MOTR, with approximately twice the inference speed. Furthermore, DecoderTracker requires significantly less training time compared to MOTR.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderMulti-Object TrackingMultiple Object TrackingObjectobject-detectionObject DetectionObject TrackingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
PatchTrack: Multiple Object Tracking Using Frame Patches
Object motion and object appearance are commonly used information in multiple object tracking (MOT) applications, either for associating detections across frames in tracking-by-detection methods or direct track predictio…
DecoderMultiple Object TrackingObjectObject TrackingTransMOT: Spatial-Temporal Graph Transformer for Multiple Object Tracking
Tracking multiple objects in videos relies on modeling the spatial-temporal interactions of the objects. In this paper, we propose a solution named TransMOT, which leverages powerful graph transformers to efficiently mod…
DecoderMulti-Object TrackingMultiple Object TrackingObject+2Joint Spatial-Temporal and Appearance Modeling with Transformer for Multiple Object Tracking
The recent trend in multiple object tracking (MOT) is heading towards leveraging deep learning to boost the tracking performance. In this paper, we propose a novel solution named TransSTAM, which leverages Transformer to…
DecoderMultiple Object TrackingObject TrackingTrackSSM: A General Motion Predictor by State-Space Model
Temporal motion modeling has always been a key component in multiple object tracking (MOT) which can ensure smooth trajectory movement and provide accurate positional information to enhance association precision. However…
DecoderMambaMulti-Object TrackingMultiple Object Tracking+4Jointly Optimizing State Operation Prediction and Value Generation for Dialogue State Tracking
We investigate the problem of multi-domain Dialogue State Tracking (DST) with open vocabulary. Existing approaches exploit BERT encoder and copy-based RNN decoder, where the encoder predicts the state operation, and the …
DecoderDialogue State TrackingMulti-domain Dialogue State Tracking