paper-with-me

홈 › Papers

Text2Control3D: Controllable 3D Avatar Generation in Neural Radiance Fields using Geometry-Guided Text-to-Image Diffusion Model

2023-09-07 · Sungwon Hwang, Junha Hyung, Jaegul Choo

Recent advances in diffusion models such as ControlNet have enabled geometrically controllable, high-fidelity text-to-image generation. However, none of them addresses the question of adding such controllability to text-to-3D generation. In response, we propose Text2Control3D, a controllable text-to-3D avatar generation method whose facial expression is controllable given a monocular video casually captured with hand-held camera. Our main strategy is to construct the 3D avatar in Neural Radiance Fields (NeRF) optimized with a set of controlled viewpoint-aware images that we generate from ControlNet, whose condition input is the depth map extracted from the input video. When generating the viewpoint-aware images, we utilize cross-reference attention to inject well-controlled, referential facial expression and appearance via cross attention. We also conduct low-pass filtering of Gaussian latent of the diffusion model in order to ameliorate the viewpoint-agnostic texture problem we observed from our empirical analysis, where the viewpoint-aware images contain identical textures on identical pixel positions that are incomprehensible in 3D. Finally, to train NeRF with the images that are viewpoint-aware yet are not strictly consistent in geometry, our approach considers per-image geometric variation as a view of deformation from a shared 3D canonical space. Consequently, we construct the 3D avatar in a canonical space of deformable NeRF by learning a set of per-image deformation via deformation field table. We demonstrate the empirical results and discuss the effectiveness of our method.

📄 PDF Abstract BibTeX arXiv:2309.03550

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationImage GenerationNeRFText to 3DText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

None 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

LegacyAvatars: Volumetric Face Avatars For Traditional Graphics Pipelines

2026-01-18 · Safa C. Medin, Gengyan Li, Ziqian Bai, Ruofei Du 외 arxiv

We introduce a novel representation for efficient classical rendering of photorealistic 3D face avatars. Leveraging recent advances in radiance fields anchored to parametric face models, our approach achieves controllabl…

Text2Avatar: Text to 3D Human Avatar Generation with Codebook-Driven Body Controllable Attribute

2024-01-01 · Chaoqun Gong, Yuqin Dai, Ronghui Li, Achun Bao 외

Generating 3D human models directly from text helps reduce the cost and time of character modeling. However, achieving multi-attribute controllable and realistic 3D human avatar generation is still challenging due to fea…

AttributeDisentanglementText to 3Dtext-to-3d-human

InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation

2024-05-24 · Yuchi Wang, Junliang Guo, Jianhong Bai, Runyi Yu 외

Recent talking avatar generation models have made strides in achieving realistic and accurate lip synchronization with the audio, but often fall short in controlling and conveying detailed expressions and emotions of the…

Hyper Diffusion Avatars: Dynamic Human Avatar Generation using Network Weight Space Diffusion

2025-09-04 · Dongliang Cao, Guoxing Sun, Marc Habermann, Florian Bernard arxiv

Creating human avatars is a highly desirable yet challenging task. Recent advancements in radiance field rendering have achieved unprecedented photorealism and real-time performance for personalized dynamic human avatars…

UNICA: A Unified Neural Framework for Controllable 3D Avatars

2026-04-03 · Jiahe Zhu, Xinyao Wang, Yiyu Zhuang, Yanwen Wang 외 arxiv

Controllable 3D human avatars have found widespread applications in 3D games, the metaverse, and AR/VR scenarios. The conventional approach to creating such a 3D avatar requires a lengthy, intricate pipeline encompassing…

Motion Planning