Weakly Supervised Monocular 3D Object Detection using Multi-View Projection and Direction Consistency
Monocular 3D object detection has become a mainstream approach in automatic driving for its easy application. A prominent advantage is that it does not need LiDAR point clouds during the inference. However, most current methods still rely on 3D point cloud data for labeling the ground truths used in the training phase. This inconsistency between the training and inference makes it hard to utilize the large-scale feedback data and increases the data collection expenses. To bridge this gap, we propose a new weakly supervised monocular 3D objection detection method, which can train the model with only 2D labels marked on images. To be specific, we explore three types of consistency in this task, i.e. the projection, multi-view and direction consistency, and design a weakly-supervised architecture based on these consistencies. Moreover, we propose a new 2D direction labeling method in this task to guide the model for accurate rotation direction prediction. Experiments show that our weakly-supervised method achieves comparable performance with some fully supervised methods. When used as a pre-training method, our model can significantly outperform the corresponding fully-supervised baseline with only 1/3 3D labels. https://github.com/weakmono3d/weakmono3d
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object DetectionMonocular 3D Object Detectionobject-detectionObject DetectionSimilar Papers 제목 키워드 기반
Weakly Supervised Monocular 3D Detection with a Single-View Image
Monocular 3D detection (M3D) aims for precise 3D object localization from a single-view image which usually involves labor-intensive annotation of 3D detection boxes. Weakly supervised M3D has recently been studied to ob…
Knowledge DistillationObject LocalizationSelf-Knowledge DistillationTransfer LearningVSRD: Instance-Aware Volumetric Silhouette Rendering for Weakly Supervised 3D Object Detection
Monocular 3D object detection poses a significant challenge in 3D scene understanding due to its inherently ill-posed nature in monocular depth estimation. Existing methods heavily rely on supervised learning using abund…
3D Object DetectionDepth EstimationMonocular 3D Object DetectionMonocular Depth Estimation+5VSRD++: Autolabeling for 3D Object Detection via Instance-Aware Volumetric Silhouette Rendering
Monocular 3D object detection is a fundamental yet challenging task in 3D scene understanding. Existing approaches heavily depend on supervised learning with extensive 3D annotations, which are often acquired from LiDAR …
Monocular 3D Object DetectionScene UnderstandingPoint CloudsWeakM3D: Towards Weakly Supervised Monocular 3D Object Detection
Monocular 3D object detection is one of the most challenging tasks in 3D scene understanding. Due to the ill-posed nature of monocular imagery, existing monocular 3D detection methods highly rely on training with the man…
3D Object DetectionMonocular 3D Object DetectionObjectobject-detection+3Weakly Supervised Training of Monocular 3D Object Detectors Using Wide Baseline Multi-view Traffic Camera Data
Accurate 7DoF prediction of vehicles at an intersection is an important task for assessing potential conflicts between road users. In principle, this could be achieved by a single camera system that is capable of detecti…
Autonomous VehiclesObjectPose Prediction