paper-with-me

Papers

MagicMan: Generative Novel View Synthesis of Humans with 3D-Aware Diffusion and Iterative Refinement

2024-08-26 · Xu He, Xiaoyu Li, Di Kang, Jiangnan Ye, Chaopeng Zhang, Liyang Chen, Xiangjun Gao, Han Zhang, Zhiyong Wu, Haolin Zhuang

Existing works in single-image human reconstruction suffer from weak generalizability due to insufficient training data or 3D inconsistencies for a lack of comprehensive multi-view knowledge. In this paper, we introduce MagicMan, a human-specific multi-view diffusion model designed to generate high-quality novel view images from a single reference image. As its core, we leverage a pre-trained 2D diffusion model as the generative prior for generalizability, with the parametric SMPL-X model as the 3D body prior to promote 3D awareness. To tackle the critical challenge of maintaining consistency while achieving dense multi-view generation for improved 3D human reconstruction, we first introduce hybrid multi-view attention to facilitate both efficient and thorough information interchange across different views. Additionally, we present a geometry-aware dual branch to perform concurrent generation in both RGB and normal domains, further enhancing consistency via geometry cues. Last but not least, to address ill-shaped issues arising from inaccurate SMPL-X estimation that conflicts with the reference image, we propose a novel iterative refinement strategy, which progressively optimizes SMPL-X accuracy while enhancing the quality and consistency of the generated multi-views. Extensive experimental results demonstrate that our method significantly outperforms existing approaches in both novel view synthesis and subsequent 3D human reconstruction tasks.

📄 PDF Abstract BibTeX arXiv:2408.14211

Code (0)

등록된 구현이 없습니다.

Tasks

3D Human ReconstructionNovel View Synthesis

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

EgoAnimate: Generating Human Animations from Egocentric top-down Views

2025-07-12 · G. Kutay Türkoglu, Julian Tanke, Iheb Belgacem, Lev Markhasin arxiv

An ideal digital telepresence experience requires accurate replication of a person's body, clothing, and movements. To capture and transfer these movements into virtual reality, the egocentric (first-person) perspective …

pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis

2020-12-02 · CVPR 2021 1 · Eric R. Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu 외

We have witnessed rapid progress on 3D-aware image synthesis, leveraging recent advances in generative visual models and neural rendering. Existing approaches however fall short in two ways: first, they may lack an under…

3D-Aware Image SynthesisImage GenerationNeural RenderingScene Generation

AvatarGen: A 3D Generative Model for Animatable Human Avatars

2022-11-26 · Jianfeng Zhang, Zihang Jiang, Dingdong Yang, Hongyi Xu 외

Unsupervised generation of 3D-aware clothed humans with various appearances and controllable geometries is important for creating virtual human avatars and other AR/VR applications. Existing methods are either limited to…

Human Animation

Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis

2022-04-13 · CVPR 2022 1 · Xuanmeng Zhang, Zhedong Zheng, Daiheng Gao, Bang Zhang 외

3D-aware image synthesis aims to generate images of objects from multiple views by learning a 3D representation. However, one key challenge remains: existing approaches lack geometry constraints, hence usually fail to ge…

3D-Aware Image Synthesis3D geometryImage Generation

Replay: Multi-modal Multi-view Acted Videos for Casual Holography

2023-07-22 · ICCV 2023 1 · Roman Shapovalov, Yanir Kleiman, Ignacio Rocco, David Novotny 외

We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as weara…

3D ReconstructionNovel View Synthesis