paper-with-me

홈 › Papers

GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human Data

2024-11-27 · Wentao Wang, Hang Ye, Fangzhou Hong, Xue Yang, Jianfu Zhang, Yizhou Wang, Ziwei Liu, Liang Pan

Given a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying body proportions captured by in-the-wild human images; b) diverse personal belongings within the shot; and c) ambiguities in human postures and inconsistency in human textures. In addition, the scarcity of high-quality human data intensifies the challenge. To address these problems, we propose a Generalizable image-to-3D huMAN reconstruction framework, dubbed GeneMAN, building upon a comprehensive multi-source collection of high-quality human data, including 3D scans, multi-view videos, single photos, and our generated synthetic human data. GeneMAN encompasses three key modules. 1) Without relying on parametric human models (e.g., SMPL), GeneMAN first trains a human-specific text-to-image diffusion model and a view-conditioned diffusion model, serving as GeneMAN 2D human prior and 3D human prior for reconstruction, respectively. 2) With the help of the pretrained human prior models, the Geometry Initialization-&-Sculpting pipeline is leveraged to recover high-quality 3D human geometry given a single image. 3) To achieve high-fidelity 3D human textures, GeneMAN employs the Multi-Space Texture Refinement pipeline, consecutively refining textures in the latent and the pixel spaces. Extensive experimental results demonstrate that GeneMAN could generate high-quality 3D human models from a single image input, outperforming prior state-of-the-art methods. Notably, GeneMAN could reveal much better generalizability in dealing with in-the-wild images, often yielding high-quality 3D human models in natural poses with common items, regardless of the body proportions in the input images.

📄 PDF Abstract BibTeX arXiv:2411.18624

Code (0)

등록된 구현이 없습니다.

Tasks

3D Human ReconstructionImage to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors

2024-06-18 · Panwang Pan, Zhuo Su, Chenguo Lin, Zhen Fan 외

Despite recent advancements in high-fidelity human reconstruction techniques, the requirements for densely captured images or time-consuming per-instance optimization significantly hinder their applications in broader sc…

Novel View Synthesis

HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D Segmentation

2025-11-01 · Panwang Pan, Tingting Shen, Chenxin Li, Yunlong Lin 외 arxiv

Recent advances in generative models have achieved high-fidelity in 3D human reconstruction, yet their utility for specific tasks (e.g., human 3D segmentation) remains constrained. We propose HumanCrafter, a unified fram…

3D Human Reconstruction

Generalizable Human Gaussians from Single-View Image

2024-06-10 · Jinnan Chen, Chen Li, Jianfeng Zhang, Lingting Zhu 외

In this work, we tackle the task of learning 3D human Gaussians from a single image, focusing on recovering detailed appearance and geometry including unobserved regions. We introduce a single-view generalizable Human Ga…

Novel View SynthesisSSIMSurface Reconstruction

NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses

2025-11-20 · Jing Wen, Alexander G. Schwing, Shenlong Wang arxiv

We tackle the task of recovering an animatable 3D human avatar from a single or a sparse set of images. For this task, beyond a set of images, many prior state-of-the-art methods use accurate "ground-truth" camera poses …

DiHuR: Diffusion-Guided Generalizable Human Reconstruction

2024-11-16 · Jinnan Chen, Chen Li, Gim Hee Lee

We introduce DiHuR, a novel Diffusion-guided model for generalizable Human 3D Reconstruction and view synthesis from sparse, minimally overlapping images. While existing generalizable human radiance fields excel at novel…

3D ReconstructionNovel View SynthesisTransfer Learning