paper-with-me

홈 › Papers

Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation

2025-09-23 · Sherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang, Yifeng Jiang, Haithem Turki, Andrea Tagliasacchi, David B. Lindell, Zan Gojcic, Sanja Fidler, Huan Ling, Jun Gao, Xuanchi Ren arxiv

The ability to generate virtual environments is crucial for applications ranging from gaming to physical AI domains such as robotics, autonomous driving, and industrial AI. Current learning-based 3D reconstruction methods rely on the availability of captured real-world multi-view data, which is not always readily available. Recent advancements in video diffusion models have shown remarkable imagination capabilities, yet their 2D nature limits the applications to simulation where a robot needs to navigate and interact with the environment. In this paper, we propose a self-distillation framework that aims to distill the implicit 3D knowledge in the video diffusion models into an explicit 3D Gaussian Splatting (3DGS) representation, eliminating the need for multi-view training data. Specifically, we augment the typical RGB decoder with a 3DGS decoder, which is supervised by the output of the RGB decoder. In this approach, the 3DGS decoder can be purely trained with synthetic data generated by video diffusion models. At inference time, our model can synthesize 3D scenes from either a text prompt or a single image for real-time rendering. Our framework further extends to dynamic 3D scene generation from a monocular input video. Experimental results show that our framework achieves state-of-the-art performance in static and dynamic 3D scene generation.

📄 PDF Abstract BibTeX arXiv:2509.19296

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving3D ReconstructionScene Generation

Similar Papers 제목 키워드 기반

Lyra 2.0: Explorable Generative 3D Worlds

2026-04-14 · Tianchang Shen, Sherwin Bahmani, Kai He, Sangeetha Grama Srinivasan 외 arxiv

Recent advances in video generation enable a new paradigm for 3D scene creation: generating camera-controlled videos that simulate scene walkthroughs, then lifting them to 3D via feed-forward reconstruction techniques. T…

Video Generation

Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction

2026-01-07 · Jiaxin Huang, Yuanbo Yang, Bangbang Yang, Lin Ma 외 arxiv

We present Gen3R, a method that bridges the strong priors of foundational reconstruction models and video diffusion models for scene-level 3D generation. We repurpose the VGGT reconstruction model to produce geometric la…

Scene Generation3D GenerationPoint Clouds

GR3EN: Generative Relighting for 3D Environments

2026-01-22 · Xiaoyan Xing, Philipp Henzler, Junhwa Hur, Runze Li 외 arxiv

We present a method for relighting 3D reconstructions of large room-scale environments. Existing solutions for 3D scene relighting often require solving under-determined or ill-conditioned inverse rendering problems, and…

Inverse Rendering3D Reconstruction

Compressing Scene Dynamics: A Generative Approach

2024-10-13 · Shanzhi Yin, Zihan Zhang, Bolin Chen, Shiqi Wang 외

This paper proposes to learn generative priors from the motion patterns instead of video contents for generative video compression. The priors are derived from small motion dynamics in common scenes such as swinging tree…

DecoderVideo Compression

ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model

2024-08-29 · Fangfu Liu, Wenqiang Sun, HanYang Wang, Yikai Wang 외

Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scen…

3D Scene Reconstruction