paper-with-me

홈 › Papers

Multi-View Depth Consistent Image Generation Using Generative AI Models: Application on Architectural Design of University Buildings

2025-03-05 · Xusheng Du, Ruihan Gui, Zhengyang Wang, Ye Zhang, Haoran Xie

In the early stages of architectural design, shoebox models are typically used as a simplified representation of building structures but require extensive operations to transform them into detailed designs. Generative artificial intelligence (AI) provides a promising solution to automate this transformation, but ensuring multi-view consistency remains a significant challenge. To solve this issue, we propose a novel three-stage consistent image generation framework using generative AI models to generate architectural designs from shoebox model representations. The proposed method enhances state-of-the-art image generation diffusion models to generate multi-view consistent architectural images. We employ ControlNet as the backbone and optimize it to accommodate multi-view inputs of architectural shoebox models captured from predefined perspectives. To ensure stylistic and structural consistency across multi-view images, we propose an image space loss module that incorporates style loss, structural loss and angle alignment loss. We then use depth estimation method to extract depth maps from the generated multi-view images. Finally, we use the paired data of the architectural images and depth maps as inputs to improve the multi-view consistency via the depth-aware 3D attention module. Experimental results demonstrate that the proposed framework can generate multi-view architectural images with consistent style and structural coherence from shoebox model inputs.

📄 PDF Abstract BibTeX arXiv:2503.03068

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationImage Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation

2024-04-04 · CVPR 2024 1 · Hanzhe Hu, Zhizhuo Zhou, Varun Jampani, Shubham Tulsiani

We present MVD-Fusion: a method for single-view 3D inference via generative modeling of multi-view-consistent RGB-D images. While recent methods pursuing 3D inference advocate learning novel-view generative models, these…

DenoisingDepth EstimationDepth Prediction

Zero-Shot Novel View and Depth Synthesis with Multi-View Geometric Diffusion

2025-01-30 · CVPR 2025 1 · Vitor Guizilini, Muhammad Zubair Irshad, Dian Chen, Greg Shakhnarovich 외

Current methods for 3D scene reconstruction from sparse posed images employ intermediate 3D representations such as neural fields, voxel grids, or 3D Gaussians, to achieve multi-view consistent scene appearance and geome…

3D Scene ReconstructionDepth EstimationNovel View Synthesis

MultiDiff: Consistent Novel View Synthesis from a Single Image

2024-06-26 · CVPR 2024 1 · Norman Müller, Katja Schwarz, Barbara Roessle, Lorenzo Porzi 외

We introduce MultiDiff, a novel approach for consistent novel view synthesis of scenes from a single RGB image. The task of synthesizing novel views from a single reference image is highly ill-posed by nature, as there e…

Image GenerationNovel View SynthesisScene Generation

MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion

2023-07-03 · NeurIPS 2023 11 · Shitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang 외

This paper introduces MVDiffusion, a simple yet effective method for generating consistent multi-view images from text prompts given pixel-to-pixel correspondences (e.g., perspective crops from a panorama or multi-view i…

Image Generation

DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware Diffusion

2024-10-31 · Weicai Ye, Chenhao Ji, Zheng Chen, Junyao Gao 외

Diffusion-based methods have achieved remarkable achievements in 2D image or 3D object generation, however, the generation of 3D scenes and even $360^{\circ}$ images remains constrained, due to the limited number of scen…

Scene Generation