paper-with-me

Papers

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving

2025-07-01 · Zeming Chen, Hang Zhao arxiv

Multi-view image generation in autonomous driving demands consistent 3D scene understanding across camera views. Most existing methods treat this problem as a 2D image set generation task, lacking explicit 3D modeling. However, we argue that a structured representation is crucial for scene generation, especially for autonomous driving applications. This paper proposes BEV-VAE for consistent and controllable view synthesis. BEV-VAE first trains a multi-view image variational autoencoder for a compact and unified BEV latent space and then generates the scene with a latent diffusion transformer. BEV-VAE supports arbitrary view generation given camera configurations, and optionally 3D layouts. Experiments on nuScenes and Argoverse 2 (AV2) show strong performance in both 3D consistent reconstruction and generation. The code is available at: https://github.com/Czm369/bev-vae.

📄 PDF Abstract BibTeX arXiv:2507.00707

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingAutonomous DrivingScene GenerationImage Generation

Similar Papers 제목 키워드 기반

Diffusion4D: Fast Spatial-temporal Consistent 4D Generation via Video Diffusion Models

2024-05-26 · Hanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang 외

The availability of large-scale multimodal datasets and advancements in diffusion models have significantly accelerated progress in 4D content generation. Most prior approaches rely on multiple image or video diffusion m…

STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians

2024-03-22 · Yifei Zeng, Yanqin Jiang, Siyu Zhu, Yuanxun Lu 외

Recent progress in pre-trained diffusion models and 3D generation have spurred interest in 4D content creation. However, achieving high-fidelity 4D generation with spatial-temporal consistency remains a challenge. In thi…

3D Generation

SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting

2024-08-25 · Wenrui Li, Fucheng Cai, Yapeng Mi, Zhe Yang 외

Text-driven 3D scene generation has seen significant advancements recently. However, most existing methods generate single-view images using generative models and then stitch them together in 3D space. This independent g…

3DGSImage GenerationScene Generation

Consolidating Attention Features for Multi-view Image Editing

2024-02-22 · Or Patashnik, Rinon Gal, Daniel Cohen-Or, Jun-Yan Zhu 외

Large-scale text-to-image models enable a wide range of image editing techniques, using text prompts or even spatial controls. However, applying these editing methods to multi-view images depicting a single scene leads t…

DrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion model

2023-10-11 · Xiaofan Li, Yifu Zhang, Xiaoqing Ye

With the increasing popularity of autonomous driving based on the powerful and unified bird's-eye-view (BEV) representation, a demand for high-quality and large-scale multi-view video data with accurate annotation is urg…

Autonomous DrivingImage GenerationVideo Generation