paper-with-me

홈 › Papers

HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion

2023-10-12 · Xian Liu, Jian Ren, Aliaksandr Siarohin, Ivan Skorokhodov, Yanyu Li, Dahua Lin, Xihui Liu, Ziwei Liu, Sergey Tulyakov

Despite significant advances in large-scale text-to-image models, achieving hyper-realistic human image generation remains a desirable yet unsolved task. Existing models like Stable Diffusion and DALL-E 2 tend to generate human images with incoherent parts or unnatural poses. To tackle these challenges, our key insight is that human image is inherently structural over multiple granularities, from the coarse-level body skeleton to fine-grained spatial geometry. Therefore, capturing such correlations between the explicit appearance and latent structure in one model is essential to generate coherent and natural human images. To this end, we propose a unified framework, HyperHuman, that generates in-the-wild human images of high realism and diverse layouts. Specifically, 1) we first build a large-scale human-centric dataset, named HumanVerse, which consists of 340M images with comprehensive annotations like human pose, depth, and surface normal. 2) Next, we propose a Latent Structural Diffusion Model that simultaneously denoises the depth and surface normal along with the synthesized RGB image. Our model enforces the joint learning of image appearance, spatial relationship, and geometry in a unified network, where each branch in the model complements to each other with both structural awareness and textural richness. 3) Finally, to further boost the visual quality, we propose a Structure-Guided Refiner to compose the predicted conditions for more detailed generation of higher resolution. Extensive experiments demonstrate that our framework yields the state-of-the-art performance, generating hyper-realistic human images under diverse scenarios. Project Page: https://snap-research.github.io/HyperHuman/

📄 PDF Abstract BibTeX arXiv:2310.08579

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Hyperrealistic neural decoding: Reconstruction of face stimuli from fMRI measurements via the GAN latent space

2021-01-01 · Thirza Dado, Yağmur Güçlütürk, Luca Ambrogioni, Gabrielle Ras 외

We introduce a new framework for hyperrealistic reconstruction of perceived naturalistic stimuli from brain recordings. To this end, we embrace the use of generative adversarial networks (GANs) at the earliest step of ou…

HyperLips: Hyper Control Lips with High Resolution Decoder for Talking Face Generation

2023-10-09 · Yaosen Chen, Yu Yao, Zhiqiang Li, Wei Wang 외

Talking face generation has a wide range of potential applications in the field of virtual digital humans. However, rendering high-fidelity facial video while ensuring lip synchronization is still a challenge for existin…

DecoderFace GenerationTalking Face Generation

Not All Latent Spaces Are Flat: Hyperbolic Concept Control

2026-03-14 · Maria Rosaria Briglia, Simone Facchiano, Paolo Cursi, Alessio Sampieri 외 arxiv

As modern text-to-image (T2I) models draw closer to synthesizing highly realistic content, the threat of unsafe content generation grows, and it becomes paramount to exercise control. Existing approaches steer these mode…

Taxonomy-aware Dynamic Motion Generation on Hyperbolic Manifolds

2025-09-25 · Luis Augenstein, Noémie Jaquier, Tamim Asfour, Leonel Rozo arxiv

Human-like motion generation for robots often draws inspiration from biomechanical studies, which often categorize complex human motions into hierarchical taxonomies. While these taxonomies provide rich structural inform…

Stylized Text-to-Motion Generation via Hypernetwork-Driven Low-Rank Adaptation

2026-05-13 · Junhyuk Jeon, Seokhyeon Hong, Junyong Noh arxiv

Text-driven motion diffusion models are capable of generating realistic human motions, but text alone often struggles to express fine-level nuances of motion, commonly referred to as style. Recent approaches have tackled…