paper-with-me

홈 › Papers

HoloDreamer: Holistic 3D Panoramic World Generation from Text Descriptions

2024-07-21 · Haiyang Zhou, Xinhua Cheng, Wangbo Yu, Yonghong Tian, Li Yuan

3D scene generation is in high demand across various domains, including virtual reality, gaming, and the film industry. Owing to the powerful generative capabilities of text-to-image diffusion models that provide reliable priors, the creation of 3D scenes using only text prompts has become viable, thereby significantly advancing researches in text-driven 3D scene generation. In order to obtain multiple-view supervision from 2D diffusion models, prevailing methods typically employ the diffusion model to generate an initial local image, followed by iteratively outpainting the local image using diffusion models to gradually generate scenes. Nevertheless, these outpainting-based approaches prone to produce global inconsistent scene generation results without high degree of completeness, restricting their broader applications. To tackle these problems, we introduce HoloDreamer, a framework that first generates high-definition panorama as a holistic initialization of the full 3D scene, then leverage 3D Gaussian Splatting (3D-GS) to quickly reconstruct the 3D scene, thereby facilitating the creation of view-consistent and fully enclosed 3D scenes. Specifically, we propose Stylized Equirectangular Panorama Generation, a pipeline that combines multiple diffusion models to enable stylized and detailed equirectangular panorama generation from complex text prompts. Subsequently, Enhanced Two-Stage Panorama Reconstruction is introduced, conducting a two-stage optimization of 3D-GS to inpaint the missing region and enhance the integrity of the scene. Comprehensive experiments demonstrated that our method outperforms prior works in terms of overall visual consistency and harmony as well as reconstruction quality and rendering robustness when generating fully enclosed scenes.

📄 PDF Abstract BibTeX arXiv:2407.15187

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

PanoWorld: Geometry-Consistent Panoramic Video World Modeling

2026-05-14 · Le Jiang, Xiangyu Bai, Bishoy Galoaa, Shayda Moezzi 외 arxiv

We present PanoWorld, a panoramic video world model that generates geometry-consistent 360$\degree$ video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do n…

Video Generation

PanoContext-Former: Panoramic Total Scene Understanding with a Transformer

2023-05-21 · CVPR 2024 1 · Yuan Dong, Chuan Fang, Liefeng Bo, Zilong Dong 외

Panoramic image enables deeper understanding and more holistic perception of $360^\circ$ surrounding environment, which can naturally encode enriched scene context information compared to standard perspective image. Prev…

3D Object Detectionobject-detectionObject DetectionScene Understanding

Review on Panoramic Imaging and Its Applications in Scene Understanding

2022-05-11 · Shaohua Gao, Kailun Yang, Hao Shi, Kaiwei Wang 외

With the rapid development of high-speed communication and artificial intelligence technologies, human perception of real-world scenes is no longer limited to the use of small Field of View (FoV) and low-dimensional scen…

Autonomous DrivingDepth EstimationImage SegmentationScene Understanding+2

PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion

2025-09-29 · Yuyang Yin, HaoXiang Guo, Fangfu Liu, Mengyu Wang 외 arxiv

Generating a complete and explorable 360-degree visual world enables a wide range of downstream applications. While prior works have advanced the field, they remain constrained by either narrow field-of-view limitations,…

Video Generation

HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation

2025-04-30 · Haiyang Zhou, Wangbo Yu, Jiawen Guan, Xinhua Cheng 외

The rapid advancement of diffusion models holds the promise of revolutionizing the application of VR and AR technologies, which typically require scene-level 4D assets for user experience. Nonetheless, existing diffusion…

Depth EstimationScene GenerationVideo Generation