Hybrid-SORT: Weak Cues Matter for Online Multi-Object Tracking
Multi-Object Tracking (MOT) aims to detect and associate all desired objects across frames. Most methods accomplish the task by explicitly or implicitly leveraging strong cues (i.e., spatial and appearance information), which exhibit powerful instance-level discrimination. However, when object occlusion and clustering occur, spatial and appearance information will become ambiguous simultaneously due to the high overlap among objects. In this paper, we demonstrate this long-standing challenge in MOT can be efficiently and effectively resolved by incorporating weak cues to compensate for strong cues. Along with velocity direction, we introduce the confidence and height state as potential weak cues. With superior performance, our method still maintains Simple, Online and Real-Time (SORT) characteristics. Also, our method shows strong generalization for diverse trackers and scenarios in a plug-and-play and training-free manner. Significant and consistent improvements are observed when applying our method to 5 different representative trackers. Further, with both strong and weak cues, our method Hybrid-SORT achieves superior performance on diverse benchmarks, including MOT17, MOT20, and especially DanceTrack where interaction and severe occlusion frequently happen with complex motions. The code and models are available at https://github.com/ymzis69/HybridSORT.
Code (2)
Tasks
Multi-Object TrackingMultiple Object TrackingObject TrackingOnline Multi-Object TrackingSimilar Papers 제목 키워드 기반
FusionSORT: Fusion Methods for Online Multi-object Visual Tracking
In this work, we investigate four different fusion methods for associating detections to tracklets in multi-object visual tracking. In addition to considering strong cues such as motion and appearance information, we als…
ObjectVisual TrackingOVG-HQ: Online Video Grounding with Hybrid-modal Queries
Video grounding (VG) task focuses on locating specific moments in a video based on a query, usually in text form. However, traditional VG struggles with some scenarios like streaming video or queries using visual cues. T…
Video GroundingUnderstanding Identity Continuity in Thermal Video through Scene-Level Consistency
Thermal pedestrian MOT remains challenging because weak appearance cues and frequent detection interruptions cause severe trajectory fragmentation. We study whether lightweight post-processing can recover identity contin…
Multi-User Remote lab: Timetable Scheduling Using Simplex Nondominated Sorting Genetic Algorithm
The scheduling of multi-user remote laboratories is modeled as a multimodal function for the proposed optimization algorithm. The hybrid optimization algorithm, hybridization of the Nelder-Mead Simplex algorithm and Non-…
SchedulingCalibrating the Nonlinear Matter Power Spectrum: Requirements for Future Weak Lensing Surveys
Uncertainties in predicting the nonlinear clustering of matter are among the most serious theoretical systematics facing the upcoming wide-field weak gravitational lensing surveys. We estimate the accuracy with which the…
Clustering