paper-with-me

Papers

WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving

2025-09-27 · Ziyue Zhu, Zhanqian Wu, Zhenxin Zhu, Lijun Zhou, Haiyang Sun, Bing Wan, Kun Ma, Guang Chen, Hangjun Ye, Jin Xie, jian Yang arxiv

Recent advances in driving-scene generation and reconstruction have demonstrated significant potential for enhancing autonomous driving systems by producing scalable and controllable training data. Existing generation methods primarily focus on synthesizing diverse and high-fidelity driving videos; however, due to limited 3D consistency and sparse viewpoint coverage, they struggle to support convenient and high-quality novel-view synthesis (NVS). Conversely, recent 3D/4D reconstruction approaches have significantly improved NVS for real-world driving scenes, yet inherently lack generative capabilities. To overcome this dilemma between scene generation and reconstruction, we propose WorldSplat, a novel feed-forward framework for 4D driving-scene generation. Our approach effectively generates consistent multi-track videos through two key steps: (i) We introduce a 4D-aware latent diffusion model integrating multi-modal information to produce pixel-aligned 4D Gaussians in a feed-forward manner. (ii) Subsequently, we refine the novel view videos rendered from these Gaussians using a enhanced video diffusion model. Extensive experiments conducted on benchmark datasets demonstrate that WorldSplat effectively generates high-fidelity, temporally and spatially consistent multi-track novel view driving videos. Project: https://wm-research.github.io/worldsplat/

📄 PDF Abstract BibTeX arXiv:2509.23402

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingScene Generation

Similar Papers 제목 키워드 기반

EVolSplat4D: Efficient Volume-based Gaussian Splatting for 4D Urban Scene Synthesis

2026-01-22 · Sheng Miao, Sijin Li, Pan Wang, Dongfeng Bai 외 arxiv

Novel view synthesis (NVS) of static and dynamic urban scenes is essential for autonomous driving simulation, yet existing methods often struggle to balance reconstruction time with quality. While state-of-the-art neural…

Novel View SynthesisAutonomous Driving

Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction

2024-12-09 · CVPR 2025 1 · Dongxu Wei, Zhiqi Li, Peidong Liu

Prior works employing pixel-based Gaussian representation have demonstrated efficacy in feed-forward sparse-view reconstruction. However, such representation necessitates cross-view overlap for accurate depth estimation,…

Autonomous DrivingDepth Estimation

DrivingForward: Feed-forward 3D Gaussian Splatting for Driving Scene Reconstruction from Flexible Surround-view Input

2024-09-19 · Qijian Tian, Xin Tan, Yuan Xie, Lizhuang Ma

We propose DrivingForward, a feed-forward Gaussian Splatting model that reconstructs driving scenes from flexible surround-view input. Driving scene images from vehicle-mounted cameras are typically sparse, with limited …

Learning Stable Canonical Worlds for Novel View Synthesis and Beyond

2026-06-22 · Xiaoyu Xu, Jian Zou, Sheyang Tang, Zhihua Wang 외 arxiv

Feed-forward Gaussian splatting (FFGS) facilitates real-time novel view synthesis, yet current methods often remain tied to view-dependent predictions. As more input views are added, they may accumulate noisy or redundan…

Semantic SegmentationNovel View Synthesis

Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views

2026-06-28 · Mijin Yoo, In Cho, Subin Jeon, Jiwoo Lee 외 hf

A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be rec…

Instance SegmentationNovel View Synthesis