paper-with-me

홈 › Papers

RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models

2024-07-09 · BoWen Zhang, Yiji Cheng, Chunyu Wang, Ting Zhang, Jiaolong Yang, Yansong Tang, Feng Zhao, Dong Chen, Baining Guo

We present RodinHD, which can generate high-fidelity 3D avatars from a portrait image. Existing methods fail to capture intricate details such as hairstyles which we tackle in this paper. We first identify an overlooked problem of catastrophic forgetting that arises when fitting triplanes sequentially on many avatars, caused by the MLP decoder sharing scheme. To overcome this issue, we raise a novel data scheduling strategy and a weight consolidation regularization term, which improves the decoder's capability of rendering sharper details. Additionally, we optimize the guiding effect of the portrait image by computing a finer-grained hierarchical representation that captures rich 2D texture cues, and injecting them to the 3D diffusion model at multiple layers via cross-attention. When trained on 46K avatars with a noise schedule optimized for triplanes, the resulting model can generate 3D avatars with notably better details than previous methods and can generalize to in-the-wild portrait input.

📄 PDF Abstract BibTeX arXiv:2407.06938

Code (1)

rodinhd/rodinhd pytorch

Tasks

DecoderScheduling

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

StyleAvatar3D: Leveraging Image-Text Diffusion Models for High-Fidelity 3D Avatar Generation

2023-05-30 · Chi Zhang, YiWen Chen, Yijun Fu, Zhenglin Zhou 외

The recent advancements in image-text diffusion models have stimulated research interest in large-scale 3D generative models. Nevertheless, the limited availability of diverse 3D resources presents significant challenges…

3D GenerationAttributeDiversityGenerative Adversarial Network

TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing

2025-11-20 · Eddie Pokming Sheung, Qihao Liu, Wufei Ma, Prakhar Kaushik 외 arxiv

With the increasing demand for 3D animation, generating high-fidelity, controllable 4D avatars from textual descriptions remains a significant challenge. Despite notable efforts in 4D generative modeling, existing method…

Computational Efficiency

AvatarStudio: High-fidelity and Animatable 3D Avatar Creation from Text

2023-11-29 · Jianfeng Zhang, Xuanmeng Zhang, Huichao Zhang, Jun Hao Liew 외

We study the problem of creating high-fidelity and animatable 3D avatars from only textual descriptions. Existing text-to-avatar methods are either limited to static avatars which cannot be animated or struggle to genera…

NeRF

GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar Reconstruction

2026-02-27 · Chao Xu, Xiaochen Zhao, Xiang Deng, Jingxiang Sun 외 arxiv

Reconstructing photorealistic and animatable 4D head avatars from a single portrait image remains a fundamental challenge in computer vision. While diffusion models have enabled remarkable progress in image and video gen…

Video Generation

Rodin: A Generative Model for Sculpting 3D Digital Avatars Using Diffusion

2022-12-12 · CVPR 2023 1 · Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu 외

This paper presents a 3D generative model that uses diffusion models to automatically generate 3D digital avatars represented as neural radiance fields. A significant challenge in generating such avatars is that the memo…

Computational Efficiency