paper-with-me

홈 › Papers

FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis

2025-03-17 · Luxi Chen, Zihan Zhou, Min Zhao, Yikai Wang, Ge Zhang, Wenhao Huang, Hao Sun, Ji-Rong Wen, Chongxuan Li

Generating flexible-view 3D scenes, including 360{\deg} rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework consisting of two key components: (1) a strong video-to-video (V2V) diffusion model to generate high-quality novel view images from incomplete input rendered from a coarse scene, and (2) a progressive expansion process to construct a complete 3D scene. In particular, leveraging an advanced pre-trained video model and accurate depth-estimated training pairs, our V2V model can generate novel views under large camera pose variations. Building upon it, FlexWorld progressively generates new 3D content and integrates it into the global scene through geometry-aware scene fusion. Extensive experiments demonstrate the effectiveness of FlexWorld in generating high-quality novel view videos and flexible-view 3D scenes from single images, achieving superior visual quality under multiple popular metrics and datasets compared to existing state-of-the-art methods. Qualitatively, we highlight that FlexWorld can generate high-fidelity scenes with flexible views like 360{\deg} rotations and zooming. Project page: https://ml-gsai.github.io/FlexWorld.

📄 PDF Abstract BibTeX arXiv:2503.13265

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Anisotropic Neural Representation Learning for High-Quality Neural Rendering

2023-11-30 · Y. Wang, J. Xu, Y. Zeng, Y. Gong

Neural radiance fields (NeRFs) have achieved impressive view synthesis results by learning an implicit volumetric representation from multi-view images. To project the implicit representation into an image, NeRF employs …

NeRFNeural RenderingRepresentation Learning

FreeSim: Toward Free-viewpoint Camera Simulation in Driving Scenes

2024-12-04 · CVPR 2025 1 · Lue Fan, Hao Zhang, Qitai Wang, Hongsheng Li 외

We propose FreeSim, a camera simulation method for autonomous driving. FreeSim emphasizes high-quality rendering from viewpoints beyond the recorded ego trajectories. In such viewpoints, previous methods have unacceptabl…

Autonomous Driving

NavCrafter: Exploring 3D Scenes from a Single Image

2026-04-03 · Hongbo Duan, Peiyu Zhuang, Yi Liu, Zhengyang Zhang 외 arxiv

Creating flexible 3D scenes from a single image is vital when direct 3D data acquisition is costly or impractical. We introduce NavCrafter, a novel framework that explores 3D scenes from a single image by synthesizing no…

3D Reconstruction

Self-Evolving 3D Scene Generation from a Single Image

2025-12-09 · Kaizhi Zheng, Yue Fan, Jing Gu, Zishuo Xu 외 arxiv

Generating high-quality, textured 3D scenes from a single image remains a fundamental challenge in vision and graphics. Recent image-to-3D generators recover reasonable geometry from single views, but their object-centri…

Scene GenerationVideo Generation3D Generation

Gaussian Scenes: Pose-Free Sparse-View Scene Reconstruction using Depth-Enhanced Diffusion Priors

2024-11-24 · Soumava Paul, Prakhar Kaushik, Alan Yuille

In this work, we introduce a generative approach for pose-free reconstruction of $360^{\circ}$ scenes from a limited number of uncalibrated 2D images. Pose-free scene reconstruction from incomplete, unposed observations …

Depth EstimationInstruction Following