Exploring Disentangled and Controllable Human Image Synthesis: From End-to-End to Stage-by-Stage
Achieving fine-grained controllability in human image synthesis is a long-standing challenge in computer vision. Existing methods primarily focus on either facial synthesis or near-frontal body generation, with limited ability to simultaneously control key factors such as viewpoint, pose, clothing, and identity in a disentangled manner. In this paper, we introduce a new disentangled and controllable human synthesis task, which explicitly separates and manipulates these four factors within a unified framework. We first develop an end-to-end generative model trained on MVHumanNet for factor disentanglement. However, the domain gap between MVHumanNet and in-the-wild data produces unsatisfactory results, motivating the exploration of virtual try-on (VTON) dataset as a potential solution. Through experiments, we observe that simply incorporating the VTON dataset as additional data to train the end-to-end model degrades performance, primarily due to the inconsistency in data forms between the two datasets, which disrupts the disentanglement process. To better leverage both datasets, we propose a stage-by-stage framework that decomposes human image generation into three sequential steps: clothed A-pose generation, back-view synthesis, and pose and view control. This structured pipeline enables better dataset utilization at different stages, significantly improving controllability and generalization, especially for in-the-wild scenarios. Extensive experiments demonstrate that our stage-by-stage approach outperforms end-to-end models in both visual fidelity and disentanglement quality, offering a scalable solution for real-world tasks. Additional demos are available on the project page: https://taited.github.io/discohuman-project/.
Code (0)
등록된 구현이 없습니다.
Tasks
DisentanglementImage GenerationVirtual Try-onMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Neural Texture Extraction and Distribution for Controllable Person Image Synthesis
We deal with the controllable person image synthesis task which aims to re-render a human from a reference image with explicit control over body pose and appearance. Observing that person images are highly structured, we…
Image GenerationSemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
With the development of neural radiance fields and generative models, numerous methods have been proposed for learning 3D human generation from 2D images. These methods allow control over the pose of the generated 3D hum…
3D-Aware Image SynthesisDisentanglementImage GenerationSuper-ResolutionHOIDiffusion: Generating Realistic 3D Hand-Object Interaction Data
3D hand-object interaction data is scarce due to the hardware constraints in scaling up the data collection process. In this paper, we propose HOIDiffusion for generating realistic and diverse 3D hand-object interaction …
6D Pose Estimation using RGBImage GenerationObjectPose EstimationDisentangled and Controllable Face Image Generation via 3D Imitative-Contrastive Learning
We propose DiscoFaceGAN, an approach for face image generation of virtual people with disentangled, precisely-controllable latent representations for identity of non-existing people, expression, pose, and illumination. W…
Contrastive LearningDisentanglementImage GenerationEncouraging Disentangled and Convex Representation with Controllable Interpolation Regularization
We focus on controllable disentangled representation learning (C-Dis-RL), where users can control the partition of the disentangled latent space to factorize dataset attributes (concepts) for downstream tasks. Two genera…
Data AugmentationDisentanglementFairnessImage Generation+1