paper-with-me

홈 › Papers

Text-Image Conditioned Diffusion for Consistent Text-to-3D Generation

2023-12-19 · Yuze He, Yushi Bai, Matthieu Lin, Jenny Sheng, Yubin Hu, Qi Wang, Yu-Hui Wen, Yong-Jin Liu

By lifting the pre-trained 2D diffusion models into Neural Radiance Fields (NeRFs), text-to-3D generation methods have made great progress. Many state-of-the-art approaches usually apply score distillation sampling (SDS) to optimize the NeRF representations, which supervises the NeRF optimization with pre-trained text-conditioned 2D diffusion models such as Imagen. However, the supervision signal provided by such pre-trained diffusion models only depends on text prompts and does not constrain the multi-view consistency. To inject the cross-view consistency into diffusion priors, some recent works finetune the 2D diffusion model with multi-view data, but still lack fine-grained view coherence. To tackle this challenge, we incorporate multi-view image conditions into the supervision signal of NeRF optimization, which explicitly enforces fine-grained view consistency. With such stronger supervision, our proposed text-to-3D method effectively mitigates the generation of floaters (due to excessive densities) and completely empty spaces (due to insufficient densities). Our quantitative evaluations on the T$^3$Bench dataset demonstrate that our method achieves state-of-the-art performance over existing text-to-3D methods. We will make the code publicly available.

📄 PDF Abstract BibTeX arXiv:2312.11774

Code (0)

등록된 구현이 없습니다.

Tasks

3D GenerationNeRFText to 3D

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

PRedItOR: Text Guided Image Editing with Diffusion Prior

2023-02-15 · Hareesh Ravi, Sachin Kelkar, Midhun Harikumar, Ajinkya Kale

Diffusion models have shown remarkable capabilities in generating high quality and creative images conditioned on text. An interesting application of such models is structure preserving text guided image editing. Existin…

Decodertext-guided-image-editing

Multi-focal Conditioned Latent Diffusion for Person Image Synthesis

2025-03-19 · CVPR 2025 1 · Jiaqi Liu, Jichao Zahng, Paolo Rota, Nicu Sebe

The Latent Diffusion Model (LDM) has demonstrated strong capabilities in high-resolution image generation and has been widely employed for Pose-Guided Person Image Synthesis (PGPIS), yielding promising results. However, …

Image Generation

Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model

2023-10-23 · Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu 외

We report Zero123++, an image-conditioned diffusion model for generating 3D-consistent multi-view images from a single input view. To take full advantage of pretrained 2D generative priors, we develop various conditionin…

High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding

2026-03-11 · Ji Woo Hong, Hee Suk Yoon, Gwanhyeong Koo, Eunseop Yoon 외 arxiv

Recent large-scale vision-language models (VLMs) have shown remarkable text-to-image generation capabilities, yet their visual fidelity remains constrained by the discrete image tokenization, which poses a major challeng…

Text-to-Image Generation

JoDiffusion: Jointly Diffusing Image with Pixel-Level Annotations for Semantic Segmentation Promotion

2025-12-15 · Haoyu Wang, Lei Zhang, Wenrui Liu, Dengyang Jiang 외 arxiv

Given the inherently costly and time-intensive nature of pixel-level annotation, the generation of synthetic datasets comprising sufficiently diverse synthetic images paired with ground-truth pixel-level annotations has …

Semantic SegmentationImage Generation