DISP6D: Disentangled Implicit Shape and Pose Learning for Scalable 6D Pose Estimation
Scalable 6D pose estimation for rigid objects from RGB images aims at handling multiple objects and generalizing to novel objects. Building on a well-known auto-encoding framework to cope with object symmetry and the lack of labeled training data, we achieve scalability by disentangling the latent representation of auto-encoder into shape and pose sub-spaces. The latent shape space models the similarity of different objects through contrastive metric learning, and the latent pose code is compared with canonical rotations for rotation retrieval. Because different object symmetries induce inconsistent latent pose spaces, we re-entangle the shape representation with canonical rotations to generate shape-dependent pose codebooks for rotation retrieval. We show state-of-the-art performance on two benchmarks containing textureless CAD objects without category and daily objects with categories respectively, and further demonstrate improved scalability by extending to a more challenging setting of daily objects across categories.
Code (1)
Tasks
6D Pose EstimationMetric LearningPose EstimationRetrievalSelf-Supervised LearningSimilar Papers 제목 키워드 기반
D$^2$IM-Net: Learning Detail Disentangled Implicit Fields from Single Images
We present the first single-view 3D reconstruction network aimed at recovering geometric details from an input image which encompass both topological shape structures and surface features. Our key idea is to train the ne…
3D ReconstructionDecoderSingle-View 3D ReconstructionD2IM-Net: Learning Detail Disentangled Implicit Fields From Single Images
We present the first single-view 3D reconstruction network aimed at recovering geometric details from an input image which encompass both topological shape structures and surface features. Our key idea is to train th…
3D ReconstructionDecoderSingle-View 3D ReconstructionNeural Template: Topology-aware Reconstruction and Disentangled Generation of 3D Meshes
This paper introduces a novel framework called DTNet for 3D mesh reconstruction and generation via Disentangled Topology. Beyond previous works, we learn a topology-aware neural template specific to each input then defor…
i3DMM: Deep Implicit 3D Morphable Model of Human Heads
We present the first deep implicit 3D morphable model (i3DMM) of full heads. Unlike earlier morphable face models it not only captures identity-specific geometry, texture, and expressions of the frontal face, but also mo…
Disentangling 3D Attributes from a Single 2D Image: Human Pose, Shape and Garment
For visual manipulation tasks, we aim to represent image content with semantically meaningful features. However, learning implicit representations from images often lacks interpretability, especially when attributes are …
3D ReconstructionDecoderDisentanglement