paper-with-me

홈 › Papers

3D Magic Mirror: Clothing Reconstruction from a Single Image via a Causal Perspective

2022-04-27 · Zhedong Zheng, Jiayin Zhu, Wei Ji, Yi Yang, Tat-Seng Chua

This research aims to study a self-supervised 3D clothing reconstruction method, which recovers the geometry shape and texture of human clothing from a single image. Compared with existing methods, we observe that three primary challenges remain: (1) 3D ground-truth meshes of clothing are usually inaccessible due to annotation difficulties and time costs; (2) Conventional template-based methods are limited to modeling non-rigid objects, e.g., handbags and dresses, which are common in fashion images; (3) The inherent ambiguity compromises the model training, such as the dilemma between a large shape with a remote camera or a small shape with a close camera. In an attempt to address the above limitations, we propose a causality-aware self-supervised learning method to adaptively reconstruct 3D non-rigid objects from 2D images without 3D annotations. In particular, to solve the inherent ambiguity among four implicit variables, i.e., camera position, shape, texture, and illumination, we introduce an explainable structural causal map (SCM) to build our model. The proposed model structure follows the spirit of the causal map, which explicitly considers the prior template in the camera estimation and shape prediction. When optimization, the causality intervention tool, i.e., two expectation-maximization loops, is deeply embedded in our algorithm to (1) disentangle four encoders and (2) facilitate the prior template. Extensive experiments on two 2D fashion benchmarks (ATR and Market-HQ) show that the proposed method could yield high-fidelity 3D reconstruction. Furthermore, we also verify the scalability of the proposed method on a fine-grained bird dataset, i.e., CUB. The code is available at https://github.com/layumi/ 3D-Magic-Mirror .

📄 PDF Abstract BibTeX arXiv:2204.13096

Code (1)

layumi/3D-Magic-Mirror 공식 구현 pytorch

Tasks

3D ReconstructionPerson Re-IdentificationSelf-Supervised LearningSingle-View 3D Reconstruction

Similar Papers 제목 키워드 기반

Magic Clothing: Controllable Garment-Driven Image Synthesis

2024-04-15 · Weifeng Chen, Tao Gu, Yuhao Xu, Chengcai Chen

We propose Magic Clothing, a latent diffusion model (LDM)-based network architecture for an unexplored garment-driven image synthesis task. Aiming at generating customized characters wearing the target garments with dive…

Image Generation

MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation

2025-09-12 · Jia Wang, Jie Hu, Xiaoqi Ma, Hanghang Ma 외 arxiv

Text-to-image (T2I) generation has achieved remarkable progress in instruction following and aesthetics. However, a persistent challenge is the prevalence of physical artifacts, such as anatomical and structural flaws, w…

Text-to-Image GenerationInstruction Following

CloTH-VTON: Clothing Three-dimensional reconstruction for Hybrid image-based Virtual Try-ON

2020-11-30 · Matiur Rahman Minar, Heejune Ahn

Virtual clothing try-on, transferring a clothing image onto a target person image, is drawing industrial and research attention. Both 2D image-based and 3D model-based methods proposed recently have their benefits and li…

Image GenerationVirtual Try-on

Magic Mirror: ID-Preserved Video Generation in Video Diffusion Transformers

2025-01-07 · Yuechen Zhang, Yaoyang Liu, Bin Xia, Bohao Peng 외

We present Magic Mirror, a framework for generating identity-preserved videos with cinematic-level quality and dynamic motion. While recent advances in video diffusion models have shown impressive capabilities in text-to…

DiversityText-to-Video GenerationVideo Generation

SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images

2026-05-27 · Sky Cen, Wufei Ma, Guofeng Zhang, Alan Yuille 외 arxiv

Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it remains unclear whether these models capture the semantic structure…

3D Shape ReconstructionSemantic correspondence