Monocular Differentiable Rendering for Self-Supervised 3D Object Detection
3D object detection from monocular images is an ill-posed problem due to the projective entanglement of depth and scale. To overcome this ambiguity, we present a novel self-supervised method for textured 3D shape reconstruction and pose estimation of rigid objects with the help of strong shape priors and 2D instance masks. Our method predicts the 3D location and meshes of each object in an image using differentiable rendering and a self-supervised objective derived from a pretrained monocular depth estimation network. We use the KITTI 3D object detection dataset to evaluate the accuracy of the method. Experiments demonstrate that we can effectively use noisy monocular depth and differentiable rendering as an alternative to expensive 3D ground-truth labels or LiDAR information.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Object Detection3D Object Detection From Monocular Images3D Shape ReconstructionDepth EstimationMonocular Depth EstimationObjectobject-detectionObject DetectionPose EstimationSimilar Papers 제목 키워드 기반
Occlusion-Aware Self-Supervised Monocular 6D Object Pose Estimation
6D object pose estimation is a fundamental yet challenging problem in computer vision. Convolutional Neural Networks (CNNs) have recently proven to be capable of predicting reliable 6D pose estimates even under monocular…
6D Pose Estimation6D Pose Estimation using RGBDomain AdaptationObject+2CPS++: Improving Class-level 6D Pose and Shape Estimation From Monocular Images With Self-Supervised Learning
Contemporary monocular 6D pose estimation methods can only cope with a handful of object instances. This naturally hampers possible applications as, for instance, robots seamlessly integrated in everyday processes necess…
6D Pose EstimationPose EstimationRetrievalSelf-Supervised LearningTowards High Fidelity Monocular Face Reconstruction with Rich Reflectance using Self-supervised Learning and Ray Tracing
Robust face reconstruction from monocular image in general lighting conditions is challenging. Methods combining deep neural network encoders with differentiable rendering have opened up the path for very fast monocular …
3D Face ReconstructionFace ReconstructionMonocular ReconstructionSelf-Supervised LearningMSDA: Monocular Self-supervised Domain Adaptation for 6D Object Pose Estimation
Acquiring labeled 6D poses from real images is an expensive and time-consuming task. Though massive amounts of synthetic RGB images are easy to obtain, the models trained on them suffer from noticeable performance degrad…
6D Pose Estimation using RGBDomain AdaptationPose EstimationPseudo LabelSHeaP: Self-Supervised Head Geometry Predictor Learned via 2D Gaussians
Accurate, real-time 3D reconstruction of human heads from monocular images and videos underlies numerous visual applications. As 3D ground truth data is hard to come by at scale, previous methods have sought to learn fro…
3D ReconstructionEmotion Classification