paper-with-me

홈 › Papers

3DDesigner: Towards Photorealistic 3D Object Generation and Editing with Text-guided Diffusion Models

2022-11-25 · Gang Li, Heliang Zheng, Chaoyue Wang, Chang Li, Changwen Zheng, DaCheng Tao

Text-guided diffusion models have shown superior performance in image/video generation and editing. While few explorations have been performed in 3D scenarios. In this paper, we discuss three fundamental and interesting problems on this topic. First, we equip text-guided diffusion models to achieve 3D-consistent generation. Specifically, we integrate a NeRF-like neural field to generate low-resolution coarse results for a given camera view. Such results can provide 3D priors as condition information for the following diffusion process. During denoising diffusion, we further enhance the 3D consistency by modeling cross-view correspondences with a novel two-stream (corresponding to two different views) asynchronous diffusion process. Second, we study 3D local editing and propose a two-step solution that can generate 360-degree manipulated results by editing an object from a single view. Step 1, we propose to perform 2D local editing by blending the predicted noises. Step 2, we conduct a noise-to-text inversion process that maps 2D blended noises into the view-independent text embedding space. Once the corresponding text embedding is obtained, 360-degree images can be generated. Last but not least, we extend our model to perform one-shot novel view synthesis by fine-tuning on a single image, firstly showing the potential of leveraging text guidance for novel view synthesis. Extensive experiments and various applications show the prowess of our 3DDesigner. The project page is available at https://3ddesigner-diffusion.github.io/.

📄 PDF Abstract BibTeX arXiv:2211.14108

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingNeRFNovel View SynthesisVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

CADDesigner: Conceptual CAD Model Generation with a General-Purpose Agent

2025-08-01 · Fengxiao Fan, Jingzhe Ni, Xiaolong Yin, Sirui Wang 외 arxiv

Computer-Aided Design (CAD) is widely used for conceptual design and parametric 3D modeling, but typically requires a high level of expertise from designers. To lower the entry barrier and facilitate early-stage CAD mode…

Code Generation

StrandDesigner: Towards Practical Strand Generation with Sketch Guidance

2025-08-03 · Na Zhang, Moran Li, Chengming Xu, Han Feng 외 arxiv

Realistic hair strand generation is crucial for applications like computer graphics and virtual reality. While diffusion models can generate hairstyles from text or images, these inputs lack precision and user-friendline…

ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks

2026-03-29 · Samin Mahdizadeh Sani, Max Ku, Nima Jamali, Matina Mahdizadeh Sani 외 arxiv

Advances in diffusion, autoregressive, and hybrid models have enabled high-quality image synthesis for tasks such as text-to-image, editing, and reference-guided composition. Yet, existing benchmarks remain limited, eith…

Image Generation

Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes

2025-04-14 · Huijie Liu, Bingcan Wang, Jie Hu, Xiaoming Wei 외

Dish images play a crucial role in the digital era, with the demand for culturally distinctive dish images continuously increasing due to the digitization of the food industry and e-commerce. In general cases, existing t…

Image GenerationLarge Language ModelText to Image GenerationText-to-Image Generation

V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties

2025-12-12 · Ye Fang, Tong Wu, Valentin Deschaintre, Duygu Ceylan 외 arxiv

Large-scale video generation models have shown remarkable potential in modeling photorealistic appearance and lighting interactions in real-world scenes. However, a closed-loop framework that jointly understands intrinsi…

Inverse RenderingVideo Generation