paper-with-me

홈 › Papers

What Makes for Text to 360-degree Panorama Generation with Stable Diffusion?

2025-05-28 · Jinhong Ni, Chang-Bin Zhang, Qiang Zhang, Jing Zhang

Recent prosperity of text-to-image diffusion models, e.g. Stable Diffusion, has stimulated research to adapt them to 360-degree panorama generation. Prior work has demonstrated the feasibility of using conventional low-rank adaptation techniques on pre-trained diffusion models to generate panoramic images. However, the substantial domain gap between perspective and panoramic images raises questions about the underlying mechanisms enabling this empirical success. We hypothesize and examine that the trainable counterparts exhibit distinct behaviors when fine-tuned on panoramic data, and such an adaptation conceals some intrinsic mechanism to leverage the prior knowledge within the pre-trained diffusion models. Our analysis reveals the following: 1) the query and key matrices in the attention modules are responsible for common information that can be shared between the panoramic and perspective domains, thus are less relevant to panorama generation; and 2) the value and output weight matrices specialize in adapting pre-trained knowledge to the panoramic domain, playing a more critical role during fine-tuning for panorama generation. We empirically verify these insights by introducing a simple framework called UniPano, with the objective of establishing an elegant baseline for future research. UniPano not only outperforms existing methods but also significantly reduces memory usage and training time compared to prior dual-branch approaches, making it scalable for end-to-end panorama generation with higher resolution. The code will be released.

📄 PDF Abstract BibTeX arXiv:2505.22129

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model

2024-01-12 · CVPR 2024 1 · Qian Wang, Weiqi Li, Chong Mou, Xinhua Cheng 외

Panorama video recently attracts more interest in both study and application, courtesy of its immersive experience. Due to the expensive cost of capturing 360-degree panoramic videos, generating desirable panorama videos…

Video Generation

A Survey on Text-Driven 360-Degree Panorama Generation

2025-02-20 · Hai Wang, Xiaoyu Xiang, Weihao Xia, Jing-Hao Xue

The advent of text-driven 360-degree panorama generation, enabling the synthesis of 360-degree panoramic images directly from textual descriptions, marks a transformative advancement in immersive visual content creation.…

Scene GenerationSurvey

Taming Stable Diffusion for Text to 360 Panorama Image Generation

2024-01-01 · CVPR 2024 1 · Cheng Zhang, Qianyi Wu, Camilo Cruz Gambardella, Xiaoshui Huang 외

Generative models e.g. Stable Diffusion have enabled the creation of photorealistic images from text prompts. Yet the generation of 360-degree panorama images from text remains a challenge particularly due to the dea…

DenoisingImage Generation

Taming Stable Diffusion for Text to 360° Panorama Image Generation

2024-04-11 · Cheng Zhang, Qianyi Wu, Camilo Cruz Gambardella, Xiaoshui Huang 외

Generative models, e.g., Stable Diffusion, have enabled the creation of photorealistic images from text prompts. Yet, the generation of 360-degree panorama images from text remains a challenge, particularly due to the de…

DenoisingImage Generation

360PanT: Training-Free Text-Driven 360-Degree Panorama-to-Panorama Translation

2024-09-12 · Hai Wang, Jing-Hao Xue

Preserving boundary continuity in the translation of 360-degree panoramas remains a significant challenge for existing text-driven image-to-image translation methods. These methods often produce visually jarring disconti…

Image-to-Image TranslationTranslation