paper-with-me

Papers

InfiniVerse: Occupancy Guided Unbounded Scene Generation for Autonomous Driving

2026-06-30 · Xiaoyu Ye, Leheng Li, Xinyu Ji, Yingjie Cai, Hongda He, Xu Yan, Guanyi Zhao, Ying-Cong Chen, Bingbing Liu, Shuguang Cui, Zhen Li arxiv

Generating realistic, controllable, and temporally coherent urban environments is a critical yet unresolved challenge in the autonomous driving community. In this paper, we introduce InfiniVerse, a unified pipeline for long-range, 2D-3D-aligned, and controllable synthesis of dynamic urban scenes from a single frame. In practice, our approach first reconstructs a 3D occupancy representation from the input multi-view frame. This representation serves as a foundation for autoregressive scene extension along arbitrary trajectories. Subsequently, a video diffusion model translates the coarse occupancy grid into realistic, spatiotemporally consistent video sequences. Moreover, we propose a hierarchical sketch-and-refine paradigm, in which the generated videos are re-projected as image-conditioned feedback to enhance the 3D occupancy representation, establishing cross-modal alignment and mutual enhancement between the visual and spatial domains. Extensive evaluations on the Waymo Open Dataset and nuScenes demonstrate that InfiniVerse achieves state-of-the-art performance, with a FID of 6.4 and FVD of 67.97, significantly outperforming existing benchmarks in both duration and stability.

📄 PDF Abstract BibTeX arXiv:2606.31109

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingScene Generation

Similar Papers 제목 키워드 기반

UniScene: Unified Occupancy-centric Driving Scene Generation

2024-12-06 · CVPR 2025 1 · Bohan Li, Jiazhe Guo, Hongsi Liu, Yingshuang Zou 외

Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to …

Autonomous DrivingScene Generation

CLONeR: Camera-Lidar Fusion for Occupancy Grid-aided Neural Representations

2022-09-02 · Alexandra Carlson, Manikandasriram Srinivasan Ramanagopal, Nathan Tseng, Matthew Johnson-Roberson 외

Recent advances in neural radiance fields (NeRFs) achieve state-of-the-art novel view synthesis and facilitate dense estimation of scene properties. However, NeRFs often fail for large, unbounded scenes that are captured…

Depth EstimationDepth PredictionNeRFNovel View Synthesis

AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond

2026-05-25 · Haiming Zhang, Junfei Zhou, Feng Jiang, Jingzhong Li 외 arxiv

Generating high-fidelity and controllable synthetic data is critical for advancing end-to-end autonomous driving, particularly for addressing the long tail of rare safety-critical scenarios. Existing occupancy-guided met…

Autonomous Driving3D ReconstructionScene GenerationVideo Generation

OccScene: Semantic Occupancy-based Cross-task Mutual Learning for 3D Scene Generation

2024-12-15 · Bohan Li, Xin Jin, Jianan Wang, Yukai Shi 외

Recent diffusion models have demonstrated remarkable performance in both 3D scene generation and perception tasks. Nevertheless, existing methods typically separate these two processes, acting as a data augmenter to gene…

MambaScene Generation

InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

2024-12-05 · Yifan Lu, Xuanchi Ren, Jiawei Yang, Tianchang Shen 외

We present InfiniCube, a scalable method for generating unbounded dynamic 3D driving scenes with high fidelity and controllability. Previous methods for scene generation either suffer from limited scales or lack geometri…

Scene Generation