Rethinking Temporal Object Detection from Robotic Perspectives
Video object detection (VID) has been vigorously studied for years but almost all literature adopts a static accuracy-based evaluation, i.e., average precision (AP). From a robotic perspective, the importance of recall continuity and localization stability is equal to that of accuracy, but the AP is insufficient to reflect detectors' performance across time. In this paper, non-reference assessments are proposed for continuity and stability based on object tracklets. These temporal evaluations can serve as supplements to static AP. Further, we develop an online tracklet refinement for improving detectors' temporal performance through short tracklet suppression, fragment filling, and temporal location fusion. In addition, we propose a small-overlap suppression to extend VID methods to single object tracking (SOT) task so that a flexible SOT-by-detection framework is then formed. Extensive experiments are conducted on ImageNet VID dataset and real-world robotic tasks, where the superiority of our proposed approaches are validated and verified. Codes will be publicly available.
Code (0)
등록된 구현이 없습니다.
Tasks
Multi-Object TrackingObjectobject-detectionObject DetectionObject TrackingVideo Object DetectionSimilar Papers 제목 키워드 기반
3D-MAN: 3D Multi-frame Attention Network for Object Detection
3D object detection is an important module in autonomous driving and robotics. However, many existing methods focus on using single frames to perform 3D detection, and do not fully utilize information from multiple frame…
3D Object DetectionAutonomous Drivingobject-detectionObject DetectionTransPillars: Coarse-to-Fine Aggregation for Multi-Frame 3D Object Detection
3D object detection using point clouds has attracted increasing attention due to its wide applications in autonomous driving and robotics. However, most existing studies focus on single point cloud frames without harness…
3D Object DetectionAutonomous DrivingObjectobject-detection+2Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
As embodied agents operate in increasingly complex environments, the ability to perceive, track, and reason about individual object instances over time becomes essential, especially in tasks requiring sequenced interacti…
Object TrackingDetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds
Existing offboard 3D detectors always follow a modular pipeline design to take advantage of unlimited sequential point clouds. We have found that the full potential of offboard 3D detectors is not explored mainly due to …
3D Multi-Object Tracking3D Object DetectionObjectobject-detection+1Rethinking the Defocus Blur Detection Problem and A Real-Time Deep DBD Model
Defocus blur detection (DBD) is a classical low level vision task. It has recently attracted attention focusing on designing complex convolutional neural networks (CNN) which make full use of both low level features and …
Data AugmentationDefocus Blur Detection