paper-with-me

홈 › Papers

Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation

2023-03-14 · Junyoung Seo, Wooseok Jang, Min-Seop Kwak, Hyeonsu Kim, Jaehoon Ko, Junho Kim, Jin-Hwa Kim, Jiyoung Lee, Seungryong Kim

Text-to-3D generation has shown rapid progress in recent days with the advent of score distillation, a methodology of using pretrained text-to-2D diffusion models to optimize neural radiance field (NeRF) in the zero-shot setting. However, the lack of 3D awareness in the 2D diffusion models destabilizes score distillation-based methods from reconstructing a plausible 3D scene. To address this issue, we propose 3DFuse, a novel framework that incorporates 3D awareness into pretrained 2D diffusion models, enhancing the robustness and 3D consistency of score distillation-based methods. We realize this by first constructing a coarse 3D structure of a given text prompt and then utilizing projected, view-specific depth map as a condition for the diffusion model. Additionally, we introduce a training strategy that enables the 2D diffusion model learns to handle the errors and sparsity within the coarse 3D structure for robust generation, as well as a method for ensuring semantic consistency throughout all viewpoints of the scene. Our framework surpasses the limitations of prior arts, and has significant implications for 3D consistent generation of 2D diffusion models.

📄 PDF Abstract BibTeX arXiv:2303.07937

Code (1)

KU-CVLAB/3DFuse 공식 구현 pytorch

Tasks

3D GenerationNeRFSingle-View 3D ReconstructionText to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models

2024-07-02 · Fei Shen, Hu Ye, Sibo Liu, Jun Zhang 외

Recent research showcases the considerable potential of conditional diffusion models for generating consistent stories. However, current methods, which predominantly generate stories in an autoregressive and excessively …

Story Visualization

ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation

2023-09-19 · Yatong Bai, Trung Dang, Dung Tran, Kazuhito Koishida 외

Diffusion models are instrumental in text-to-audio (TTA) generation. Unfortunately, they suffer from slow inference due to an excessive number of queries to the underlying denoising network per generation. To address thi…

AudioCapsAudio GenerationDenoisingDiversity

Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models

2024-09-11 · Haibo Yang, Yang Chen, Yingwei Pan, Ting Yao 외

Despite having tremendous progress in image-to-3D generation, existing methods still struggle to produce multi-view consistent images with high-resolution textures in detail, especially in the paradigm of 2D diffusion th…

3D Generation3D ReconstructionImage GenerationImage to 3D+2

Retrieval-Augmented Score Distillation for Text-to-3D Generation

2024-02-05 · Junyoung Seo, Susung Hong, Wooseok Jang, Inès Hyeonsu Kim 외

Text-to-3D generation has achieved significant success by incorporating powerful 2D diffusion models, but insufficient 3D prior knowledge also leads to the inconsistency of 3D geometry. Recently, since large-scale multi-…

3D Generation3D geometryDiversityRetrieval+1

Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion

2024-11-15 · Haoran Wei, Wencheng Han, Xingping Dong, Jianbing Shen

Recent diffusion-based Single-image 3D portrait generation methods typically employ 2D diffusion models to provide multi-view knowledge, which is then distilled into 3D representations. However, these methods usually str…

AttributeTexture Synthesis