Autolabeling 3D Objects with Differentiable Rendering of SDF Shape Priors
We present an automatic annotation pipeline to recover 9D cuboids and 3D shapes from pre-trained off-the-shelf 2D detectors and sparse LIDAR data. Our autolabeling method solves an ill-posed inverse problem by considering learned shape priors and optimizing geometric and physical parameters. To address this challenging problem, we apply a novel differentiable shape renderer to signed distance fields (SDF), leveraged together with normalized object coordinate spaces (NOCS). Initially trained on synthetic data to predict shape and coordinates, our method uses these predictions for projective and geometric alignment over real samples. Moreover, we also propose a curriculum learning strategy, iteratively retraining on samples of increasing difficulty in subsequent self-improving annotation rounds. Our experiments on the KITTI3D dataset show that we can recover a substantial amount of accurate cuboids, and that these autolabels can be used to train 3D vehicle detectors with state-of-the-art results.
Code (1)
Tasks
Weakly Supervised 3D DetectionSimilar Papers 제목 키워드 기반
Monocular Differentiable Rendering for Self-Supervised 3D Object Detection
3D object detection from monocular images is an ill-posed problem due to the projective entanglement of depth and scale. To overcome this ambiguity, we present a novel self-supervised method for textured 3D shape reconst…
3D Object Detection3D Object Detection From Monocular Images3D Shape ReconstructionDepth Estimation+5Unified Shape and SVBRDF Recovery using Differentiable Monte Carlo Rendering
Reconstructing the shape and appearance of real-world objects using measured 2D images has been a long-standing problem in computer vision. In this paper, we introduce a new analysis-by-synthesis technique capable of pro…
GPUVSRD++: Autolabeling for 3D Object Detection via Instance-Aware Volumetric Silhouette Rendering
Monocular 3D object detection is a fundamental yet challenging task in 3D scene understanding. Existing approaches heavily depend on supervised learning with extensive 3D annotations, which are often acquired from LiDAR …
Monocular 3D Object DetectionScene UnderstandingPoint CloudsWeakly Supervised Learning of Multi-Object 3D Scene Decompositions Using Deep Shape Priors
Representing scenes at the granularity of objects is a prerequisite for scene understanding and decision making. We propose PriSMONet, a novel approach based on Prior Shape knowledge for learning Multi-Object 3D scene de…
Decision MakingScene UnderstandingWeakly-supervised LearningS2P3: Self-Supervised Polarimetric Pose Prediction
This paper proposes the first self-supervised 6D object pose prediction from multimodal RGB+polarimetric images. The novel training paradigm comprises 1) a physical model to extract geometric information of polarized lig…
Knowledge DistillationPose PredictionPrediction