paper-with-me

홈 › Papers

ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation

2025-04-03 · Yuan Zhou, Shilong Jin, Litao Hua, Wanjun Lv, Haoran Duan, Jungong Han

Recent advances in zero-shot text-to-3D generation have revolutionized 3D content creation by enabling direct synthesis from textual descriptions. While state-of-the-art methods leverage 3D Gaussian Splatting with score distillation to enhance multi-view rendering through pre-trained text-to-image (T2I) models, they suffer from inherent view biases in T2I priors. These biases lead to inconsistent 3D generation, particularly manifesting as the multi-face Janus problem, where objects exhibit conflicting features across views. To address this fundamental challenge, we propose ConsDreamer, a novel framework that mitigates view bias by refining both the conditional and unconditional terms in the score distillation process: (1) a View Disentanglement Module (VDM) that eliminates viewpoint biases in conditional prompts by decoupling irrelevant view components and injecting precise camera parameters; and (2) a similarity-based partial order loss that enforces geometric consistency in the unconditional term by aligning cosine similarities with azimuth relationships. Extensive experiments demonstrate that ConsDreamer effectively mitigates the multi-face Janus problem in text-to-3D generation, outperforming existing methods in both visual quality and consistency.

📄 PDF Abstract BibTeX arXiv:2504.02316

Code (1)

GAInuist/ConsDreamer 공식 구현 jax

Tasks

3D GenerationText to 3D

Similar Papers 제목 키워드 기반

T3D: Advancing 3D Medical Vision-Language Pre-training by Learning Multi-View Visual Consistency

2023-12-03 · Che Liu, Cheng Ouyang, Yinda Chen, Cesar César Quilodrán-Casas 외

While 3D visual self-supervised learning (vSSL) shows promising results in capturing visual representations, it overlooks the clinical knowledge from radiology reports. Meanwhile, 3D medical vision-language pre-training …

Clinical KnowledgeContrastive LearningCross-Modal RetrievalImage Restoration+5

MEt3R: Measuring Multi-View Consistency in Generated Images

2025-01-10 · CVPR 2025 1 · Mohammad Asim, Christopher Wewer, Thomas Wimmer, Bernt Schiele 외

We introduce MEt3R, a metric for multi-view consistency in generated images. Large-scale generative models for multi-view image generation are rapidly advancing the field of 3D inference from sparse observations. However…

Image GenerationVideo Generation

Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement

2025-12-09 · Xinyue Liang, Zhinyuan Ma, Lingchen Sun, Yanjun Guo 외 arxiv

Although recent 3D-native generators have made great progress in synthesizing reliable geometry, they still fall short in achieving realistic appearances. A key obstacle lies in the lack of diverse and high-quality real-…

3D Generation

Ctrl123: Consistent Novel View Synthesis via Closed-Loop Transcription

2024-03-16 · Hongxiang Zhao, Xili Dai, Jianan Wang, Shengbang Tong 외

Large image diffusion models have demonstrated zero-shot capability in novel view synthesis (NVS). However, existing diffusion-based NVS methods struggle to generate novel views that are accurately consistent with the co…

3D ReconstructionNovel View Synthesis

Survey on Monocular Metric Depth Estimation

2025-01-21 · Jiuling Zhang

Monocular Depth Estimation (MDE) is fundamental to computer vision, enabling spatial understanding, 3D reconstruction, and autonomous driving. Deep learning-based MDE predicts relative depth from a single image, but the …

3D ReconstructionAutonomous DrivingData AugmentationDepth Estimation+4