paper-with-me

Papers

DiffHuman: Probabilistic Photorealistic 3D Reconstruction of Humans

2024-01-01 · CVPR 2024 1 · Akash Sengupta, Thiemo Alldieck, Nikos Kolotouros, Enric Corona, Andrei Zanfir, Cristian Sminchisescu

We present DiffHuman a probabilistic method for photorealistic 3D human reconstruction from a single RGB image. Despite the ill-posed nature of this problem most methods are deterministic and output a single solution often resulting in a lack of geometric detail and blurriness in unseen or uncertain regions. In contrast DiffHuman predicts a probability distribution over 3D reconstructions conditioned on an input 2D image which allows us to sample multiple detailed 3D avatars that are consistent with the image. DiffHuman is implemented as a conditional diffusion model that denoises pixel-aligned 2D observations of an underlying 3D shape representation. During inference we may sample 3D avatars by iteratively denoising 2D renders of the predicted 3D representation. Furthermore we introduce a generator neural network that approximates rendering with considerably reduced runtime (55x speed up) resulting in a novel dual-branch diffusion framework. Our experiments show that DiffHuman can produce diverse and detailed reconstructions for the parts of the person that are unseen or uncertain in the input image while remaining competitive with the state-of-the-art when reconstructing visible surfaces.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

3D Human Reconstruction3D Reconstruction3D Shape RepresentationDenoising

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing

2022-04-19 · CVPR 2022 1 · Thiemo Alldieck, Mihai Zanfir, Cristian Sminchisescu

We present PHORHUM, a novel, end-to-end trainable, deep neural network methodology for photorealistic 3D human reconstruction given just a monocular RGB image. Our pixel-aligned method estimates detailed 3D geometry and,…

3D geometry3D Human Reconstruction3D Reconstruction

GUSH3R: Everyone Everywhere All at Once as Gaussians

2026-07-06 · Keito Abe, Kaede Shiohara, Takashi Otonari, Toshihiko Yamasaki arxiv

Reconstructing dynamic human-scene environments from monocular videos is a challenging problem that requires jointly modeling scene geometry, camera motion, and non-rigid human dynamics while enabling photorealistic rend…

Novel View Synthesis3D ReconstructionPoint Clouds

HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors

2024-06-18 · Panwang Pan, Zhuo Su, Chenguo Lin, Zhen Fan 외

Despite recent advancements in high-fidelity human reconstruction techniques, the requirements for densely captured images or time-consuming per-instance optimization significantly hinder their applications in broader sc…

Novel View Synthesis

IDOL: Instant Photorealistic 3D Human Creation from a Single Image

2024-12-19 · CVPR 2025 1 · Yiyu Zhuang, Jiaxi Lv, Hao Wen, Qing Shuai 외

Creating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve …

GPU

From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations

2024-01-03 · CVPR 2024 1 · Evonne Ng, Javier Romero, Timur Bagautdinov, Shaojie Bai 외

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we output multiple possibilities of gestural mot…

DiversityQuantization