Self-Supervised 3D Monocular Object Detection by Recycling Bounding Boxes
Modern object detection architectures are moving towards employing self-supervised learning (SSL) to improve performance detection with related pretext tasks. Pretext tasks for monocular 3D object detection have not yet been explored yet in literature. The paper studies the application of established self-supervised bounding box recycling by labeling random windows as the pretext task. The classifier head of the 3D detector is trained to classify random windows containing different proportions of the ground truth objects, thus handling the foreground-background imbalance. We evaluate the pretext task using the RTM3D detection model as baseline, with and without the application of data augmentation. We demonstrate improvements of between 2-3 % in mAP 3D and 0.9-1.5 % BEV scores using SSL over the baseline scores. We propose the inverse class frequency re-weighted (ICFW) mAP score that highlights improvements in detection for low frequency classes in a class imbalanced dataset with long tails. We demonstrate improvements in ICFW both mAP 3D and BEV scores to take into account the class imbalance in the KITTI validation dataset. We see 4-5 % increase in ICFW metric with the pretext task.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object DetectionData AugmentationMonocular 3D Object Detectionobject-detectionObject DetectionSelf-Supervised LearningSimilar Papers 제목 키워드 기반
Multi-Task Self-Supervised Object Detection via Recycling of Bounding Box Annotations
In spite of recent enormous success of deep convolutional networks in object detection, they require a large amount of bounding box annotations, which are often time-consuming and error-prone to obtain. To make better us…
Multi-Task LearningNovel Object DetectionObjectobject-detection+3Monocular Differentiable Rendering for Self-Supervised 3D Object Detection
3D object detection from monocular images is an ill-posed problem due to the projective entanglement of depth and scale. To overcome this ambiguity, we present a novel self-supervised method for textured 3D shape reconst…
3D Object Detection3D Object Detection From Monocular Images3D Shape ReconstructionDepth Estimation+5Time-to-Label: Temporal Consistency for Self-Supervised Monocular 3D Object Detection
Monocular 3D object detection continues to attract attention due to the cost benefits and wider availability of RGB cameras. Despite the recent advances and the ability to acquire data at scale, annotation cost and compl…
3D Object DetectionDepth EstimationMonocular 3D Object DetectionObject+23D Object Aided Self-Supervised Monocular Depth Estimation
Monocular depth estimation has been actively studied in fields such as robot vision, autonomous driving, and 3D scene understanding. Given a sequence of color images, unsupervised learning methods based on the framework …
3D Object DetectionAutonomous DrivingDepth EstimationMonocular 3D Object Detection+5View-to-Label: Multi-View Consistency for Self-Supervised 3D Object Detection
For autonomous vehicles, driving safely is highly dependent on the capability to correctly perceive the environment in 3D space, hence the task of 3D object detection represents a fundamental aspect of perception. While …
3D Object DetectionAutonomous Vehiclesobject-detectionObject Detection