Disentangling 3D Attributes from a Single 2D Image: Human Pose, Shape and Garment
For visual manipulation tasks, we aim to represent image content with semantically meaningful features. However, learning implicit representations from images often lacks interpretability, especially when attributes are intertwined. We focus on the challenging task of extracting disentangled 3D attributes only from 2D image data. Specifically, we focus on human appearance and learn implicit pose, shape and garment representations of dressed humans from RGB images. Our method learns an embedding with disentangled latent representations of these three image properties and enables meaningful re-assembling of features and property control through a 2D-to-3D encoder-decoder structure. The 3D model is inferred solely from the feature map in the learned embedding space. To the best of our knowledge, our method is the first to achieve cross-domain disentanglement for this highly under-constrained problem. We qualitatively and quantitatively demonstrate our framework's ability to transfer pose, shape, and garments in 3D reconstruction on virtual data and show how an implicit shape loss can benefit the model's ability to recover fine-grained reconstruction details.
Code (0)
등록된 구현이 없습니다.
Tasks
3D ReconstructionDecoderDisentanglementSimilar Papers 제목 키워드 기반
Scaling-up Disentanglement for Image Translation
Image translation methods typically aim to manipulate a set of labeled attributes (given as supervision at training time e.g. domain label) while leaving the unlabeled attributes intact. Current methods achieve either: (…
DisentanglementDiversityTranslationMAC: A Benchmark for Multiple Attributes Compositional Zero-Shot Learning
Compositional Zero-Shot Learning (CZSL) aims to learn semantic primitives (attributes and objects) from seen compositions and recognize unseen attribute-object compositions. Existing CZSL datasets focus on single attribu…
AttributeCompositional Zero-Shot LearningZero-Shot Learning3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow
Reconstructing 3D shape from a single 2D image is a challenging task, which needs to estimate the detailed 3D structures based on the semantic attributes from 2D image. So far, most of the previous methods still struggle…
3D Reconstruction3D Shape ReconstructionAttributeDecoderAttribute-Driven Feature Disentangling and Temporal Aggregation for Video Person Re-Identification
Video-based person re-identification plays an important role in surveillance video analysis, expanding image-based methods by learning features of multiple frames. Most existing methods fuse features by temporal average-…
AttributePerson Re-IdentificationVideo-Based Person Re-IdentificationExploiting video sequences for unsupervised disentangling in generative adversarial networks
In this work we present an adversarial training algorithm that exploits correlations in video to learn --without supervision-- an image generator model with a disentangled latent space. The proposed methodology requires …