paper-with-me

홈 › Papers

ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

2023-05-25 · NeurIPS 2023 11 · Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, Jun Zhu

Score distillation sampling (SDS) has shown great promise in text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models, but suffers from over-saturation, over-smoothing, and low-diversity problems. In this work, we propose to model the 3D parameter as a random variable instead of a constant as in SDS and present variational score distillation (VSD), a principled particle-based variational framework to explain and address the aforementioned issues in text-to-3D generation. We show that SDS is a special case of VSD and leads to poor samples with both small and large CFG weights. In comparison, VSD works well with various CFG weights as ancestral sampling from diffusion models and simultaneously improves the diversity and sample quality with a common CFG weight (i.e., $7.5$). We further present various improvements in the design space for text-to-3D such as distillation time schedule and density initialization, which are orthogonal to the distillation algorithm yet not well explored. Our overall approach, dubbed ProlificDreamer, can generate high rendering resolution (i.e., $512\times512$) and high-fidelity NeRF with rich structure and complex effects (e.g., smoke and drops). Further, initialized from NeRF, meshes fine-tuned by VSD are meticulously detailed and photo-realistic. Project page and codes: https://ml.cs.tsinghua.edu.cn/prolificdreamer/

📄 PDF Abstract BibTeX arXiv:2305.16213

Code (2)

threestudio-project/threestudio 공식 구현 jax
yuanzhi-zhu/prolific_dreamer2d pytorch

Tasks

3D GenerationDiversityNeRFText to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Semi-supervised FusedGAN for Conditional Image Generation

2018-01-17 · ECCV 2018 9 · Navaneeth Bodla, Gang Hua, Rama Chellappa

We present FusedGAN, a deep network for conditional image synthesis with controllable sampling of diverse images. Fidelity, diversity and controllable sampling are the main quality measures of a good image generation mod…

AttributeConditional Image GenerationDiversityFace Generation+2

CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language

2022-11-02 · CVPR 2023 1 · Aditya Sanghi, Rao Fu, Vivian Liu, Karl Willis 외

Recent works have demonstrated that natural language can be used to generate and edit 3D shapes. However, these methods generate shapes with limited fidelity and diversity. We introduce CLIP-Sculptor, a method to address…

DiversityImage GenerationText to 3DText-to-Shape Generation

RefAdGen: High-Fidelity Advertising Image Generation

2025-08-12 · Yiyun Chen, Weikai Yang arxiv

The rapid advancement of Artificial Intelligence Generated Content (AIGC) techniques has unlocked opportunities in generating diverse and compelling advertising images based on referenced product images and textual scene…

Data AugmentationImage Generation

ComFusion: Personalized Subject Generation in Multiple Specific Scenes From Single Image

2024-02-19 · Yan Hong, Jianfu Zhang

Recent advancements in personalizing text-to-image (T2I) diffusion models have shown the capability to generate images based on personalized visual concepts using a limited number of user-provided examples. However, thes…

DanceMosaic: High-Fidelity Dance Generation with Multimodal Editability

2025-04-06 · Foram Niravbhai Shah, Parshwa Shah, Muhammad Usama Saleem, Ekkasit Pinyoanuntapong 외

Recent advances in dance generation have enabled automatic synthesis of 3D dance motions. However, existing methods still struggle to produce high-fidelity dance sequences that simultaneously deliver exceptional realism,…

Motion GenerationMotion Synthesis