paper-with-me

홈 › Papers

4K4DGen: Panoramic 4D Generation at 4K Resolution

2024-06-19 · Renjie Li, Panwang Pan, Bangbang Yang, Dejia Xu, Shijie Zhou, Xuanyang Zhang, Zeming Li, Achuta Kadambi, Zhangyang Wang, Zhengzhong Tu, Zhiwen Fan

The blooming of virtual reality and augmented reality (VR/AR) technologies has driven an increasing demand for the creation of high-quality, immersive, and dynamic environments. However, existing generative techniques either focus solely on dynamic objects or perform outpainting from a single perspective image, failing to meet the requirements of VR/AR applications that need free-viewpoint, 360$^{\circ}$ virtual views where users can move in all directions. In this work, we tackle the challenging task of elevating a single panorama to an immersive 4D experience. For the first time, we demonstrate the capability to generate omnidirectional dynamic scenes with 360$^{\circ}$ views at 4K (4096 $\times$ 2048) resolution, thereby providing an immersive user experience. Our method introduces a pipeline that facilitates natural scene animations and optimizes a set of dynamic Gaussians using efficient splatting techniques for real-time exploration. To overcome the lack of scene-scale annotated 4D data and models, especially in panoramic formats, we propose a novel \textbf{Panoramic Denoiser} that adapts generic 2D diffusion priors to animate consistently in 360$^{\circ}$ images, transforming them into panoramic videos with dynamic scenes at targeted regions. Subsequently, we propose \textbf{Dynamic Panoramic Lifting} to elevate the panoramic video into a 4D immersive environment while preserving spatial and temporal consistency. By transferring prior knowledge from 2D models in the perspective domain to the panoramic domain and the 4D lifting with spatial appearance and geometry regularization, we achieve high-quality Panorama-to-4D generation at a resolution of 4K for the first time.

📄 PDF Abstract BibTeX arXiv:2406.13527

Code (0)

등록된 구현이 없습니다.

Tasks

4k

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction Priors

2026-02-05 · Jingdong Zhang, Xiaohang Zhan, Lingzhi Zhang, Yizhou Wang 외 arxiv

Comprehensive panoramic scene understanding is critical for immersive applications, yet it remains challenging due to the scarcity of high-resolution, multi-task annotations. While perspective foundation models have achi…

Scene Understanding

DynamicScaler: Seamless and Scalable Video Generation for Panoramic Scenes

2024-12-15 · CVPR 2025 1 · Jinxiu Liu, Shaoheng Lin, Yinxiao Li, Ming-Hsuan Yang

The increasing demand for immersive AR/VR applications and spatial intelligence has heightened the need to generate high-quality scene-level and 360{\deg} panoramic video. However, most video diffusion models are constra…

DenoisingVideo Generation

Towards Realistic Data Generation for Real-World Super-Resolution

2024-06-11 · Long Peng, Wenbo Li, Renjing Pei, Jingjing Ren 외

Existing image super-resolution (SR) techniques often fail to generalize effectively in complex real-world settings due to the significant divergence between training data and practical scenarios. To address this challen…

Image Super-ResolutionSuper-Resolution

Multi-Scale Diffusion: Enhancing Spatial Layout in High-Resolution Panoramic Image Generation

2024-10-24 · XiaoYu Zhang, Teng Zhou, Xinlong Zhang, Jia Wei 외

Diffusion models have recently gained recognition for generating diverse and high-quality content, especially in the domain of image synthesis. These models excel not only in creating fixed-size images but also in produc…

Image Generation

4DGen: Grounded 4D Content Generation with Spatial-temporal Consistency

2023-12-28 · Yuyang Yin, Dejia Xu, Zhangyang Wang, Yao Zhao 외

Aided by text-to-image and text-to-video diffusion models, existing 4D content creation pipelines utilize score distillation sampling to optimize the entire dynamic 3D scene. However, as these pipelines generate 4D conte…

Motion GenerationPrompt Engineering