paper-with-me

홈 › Papers

IDOL: Instant Photorealistic 3D Human Creation from a Single Image

2024-12-19 · CVPR 2025 1 · Yiyu Zhuang, Jiaxi Lv, Hao Wen, Qing Shuai, Ailing Zeng, Hao Zhu, Shifeng Chen, Yujiu Yang, Xun Cao, Wei Liu

Creating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the perspectives of dataset, model, and representation. First, we introduce a large-scale HUman-centric GEnerated dataset, HuGe100K, consisting of 100K diverse, photorealistic sets of human images. Each set contains 24-view frames in specific human poses, generated using a pose-controllable image-to-multi-view model. Next, leveraging the diversity in views, poses, and appearances within HuGe100K, we develop a scalable feed-forward transformer model to predict a 3D human Gaussian representation in a uniform space from a given human image. This model is trained to disentangle human pose, body shape, clothing geometry, and texture. The estimated Gaussians can be animated without post-processing. We conduct comprehensive experiments to validate the effectiveness of the proposed dataset and method. Our model demonstrates the ability to efficiently reconstruct photorealistic humans at 1K resolution from a single input image using a single GPU instantly. Additionally, it seamlessly supports various applications, as well as shape and texture editing tasks. Project page: https://yiyuzhuang.github.io/IDOL/.

📄 PDF Abstract BibTeX arXiv:2412.14963

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Seeing Faces in Things: A Model and Dataset for Pareidolia

2024-09-24 · Mark Hamilton, Simon Stent, Vasha DuTell, Anne Harrington 외

The human visual system is well-tuned to detect faces of all shapes and sizes. While this brings obvious survival advantages, such as a better chance of spotting unknown predators in the bush, it also leads to spurious f…

Morphable Diffusion: 3D-Consistent Diffusion for Single-image Avatar Creation

2024-01-09 · CVPR 2024 1 · Xiyi Chen, Marko Mihajlovic, Shaofei Wang, Sergey Prokudin 외

Recent advances in generative diffusion models have enabled the previously unfeasible capability of generating 3D assets from a single input image or a text prompt. In this work, we aim to enhance the quality and functio…

Novel View Synthesis

Visual Retrieval-Augmented Generation for Silhouette-Guided Animal Art

2026-06-16 · Quoc-Duy Tran, Anh-Tuan Vo, Trung-Nghia Le arxiv

Generative AI has advanced the ability to render photorealistic or artistic images, yet it remains limited in a key aspect of human creativity: interpreting ambiguous shapes. This phenomenon, rooted in pareidolia, allows…

FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait Image

2026-06-23 · Kim Youwang, Zhengyu Yang, Liuhao Ge, Yu Rong 외 arxiv

We introduce FiCA, a Feed-forward, instant Gaussian Codec Avatar generation pipeline that creates lifelike avatars from a single portrait image. Generating a photorealistic and drivable avatar from just a single image is…

Diamonds in the Sky: Pareidolic Animals in Clouds

2026-05-31 · Miriam Horovicz, Yacov Hel-Or, Yael Moses arxiv

People often see animal shapes in clouds, a phenomenon known as pareidolia. We propose an AI-based method that aims to predict which animals people are likely to perceive in clouds, even though state-of-the-art recogniti…