paper-with-me

Papers

StructLDM: Structured Latent Diffusion for 3D Human Generation

2024-04-01 · Tao Hu, Fangzhou Hong, Ziwei Liu

Recent 3D human generative models have achieved remarkable progress by learning 3D-aware GANs from 2D images. However, existing 3D human generative methods model humans in a compact 1D latent space, ignoring the articulated structure and semantics of human body topology. In this paper, we explore more expressive and higher-dimensional latent space for 3D human modeling and propose StructLDM, a diffusion-based unconditional 3D human generative model, which is learned from 2D images. StructLDM solves the challenges imposed due to the high-dimensional growth of latent space with three key designs: 1) A semantic structured latent space defined on the dense surface manifold of a statistical human body template. 2) A structured 3D-aware auto-decoder that factorizes the global latent space into several semantic body parts parameterized by a set of conditional structured local NeRFs anchored to the body template, which embeds the properties learned from the 2D training data and can be decoded to render view-consistent humans under different poses and clothing styles. 3) A structured latent diffusion model for generative human appearance sampling. Extensive experiments validate StructLDM's state-of-the-art generation performance and illustrate the expressiveness of the structured latent space over the well-adopted 1D latent space. Notably, StructLDM enables different levels of controllable 3D human generation and editing, including pose/view/shape control, and high-level tasks including compositional generations, part-aware clothing editing, 3D virtual try-on, etc. Our project page is at: https://taohuumd.github.io/projects/StructLDM/.

📄 PDF Abstract BibTeX arXiv:2404.01241

Code (0)

등록된 구현이 없습니다.

Tasks

Virtual Try-on

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Latent Diffusion Model Diffusion models applied to latent spaces, which are normally built with (Variational) Autoencoders.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space

2025-08-01 · Junyu Chen, Dongyun Zou, Wenkun He, Junsong Chen 외 arxiv

We present DC-AE 1.5, a new family of deep compression autoencoders for high-resolution diffusion models. Increasing the autoencoder's latent channel number is a highly effective approach for improving its reconstruction…

Image Generation

TeRA: Rethinking Text-guided Realistic 3D Avatar Generation

2025-09-02 · Yanwen Wang, Yiyu Zhuang, Jiawei Zhang, Li Wang 외 arxiv

In this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach employs a two…

Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion

2026-05-06 · Yiran Qiao, Yiren Lu, Yunlai Zhou, Disheng Liu 외 arxiv

3D asset generation plays a pivotal role in fields such as gaming and virtual reality, enabling the rapid synthesis of high-fidelity 3D objects from a single or multiple images. Building on this capability, enabling styl…

Style Transfer3D Generation

Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation

2026-02-24 · Zichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu 외 arxiv

Generating realistic 3D Human-Human Interaction (HHI) requires coherent modeling of the physical plausibility of the agents and their interaction semantics. Existing methods compress all motion information into a single …

Computational EfficiencyContrastive Learning

Contrastive Diffusion Alignment: Learning Structured Latents for Controllable Generation

2025-10-16 · Ruchi Sandilya, Sumaira Perez, Charles Lynch, Lindsay Victoria 외 arxiv

Diffusion models excel at generation, but their latent spaces are high dimensional and not explicitly organized for interpretation or control. We introduce ConDA (Contrastive Diffusion Alignment), a plug-and-play geometr…

Contrastive Learning