DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation
Recent advancements in generative models have provided promising solutions for synthesizing realistic driving videos, which are crucial for training autonomous driving perception models. However, existing approaches often struggle with multi-view video generation due to the challenges of integrating 3D information while maintaining spatial-temporal consistency and effectively learning from a unified model. We propose DriveScape, an end-to-end framework for multi-view, 3D condition-guided video generation, capable of producing 1024 x 576 high-resolution videos at 10Hz. Unlike other methods limited to 2Hz due to the 3D box annotation frame rate, DriveScape overcomes this with its ability to operate under sparse conditions. Our Bi-Directional Modulated Transformer (BiMot) ensures precise alignment of 3D structural information, maintaining spatial-temporal consistency. DriveScape excels in video generation performance, achieving state-of-the-art results on the nuScenes dataset with an FID score of 8.34 and an FVD score of 76.39. Our project homepage: https://metadrivescape.github.io/papers_project/drivescapev1/index.html
Code (0)
등록된 구현이 없습니다.
Tasks
Autonomous DrivingVideo GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
DriveScape: High-Resolution Driving Video Generation by Multi-View Feature Fusion
Recent advancements in generative models offer promising solutions for synthesizing realistic driving videos, aiding in training autonomous driving perception models. However, existing methods often struggle with hig…
Autonomous DrivingDenoisingVideo GenerationMyGo: Consistent and Controllable Multi-View Driving Video Generation with Camera Control
High-quality driving video generation is crucial for providing training data for autonomous driving models. However, current generative models rarely focus on enhancing camera motion control under multi-view tasks, which…
Autonomous DrivingVideo GenerationSceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination
SceneHI is a framework that lifts high-resolution, illumination-aware priors from 2D diffusion models to perform 3D texture synthesis. It is the first to demonstrate that high-resolution textures, previously limited to 2…
DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model
We present DreamForge-World 0.1 Preview, a preview foundational world model for real-time interactive world simulation. The system adapts the LongLive 1 autoregressive video stack, itself derived from Wan2.1-T2V-1.3B, wi…
GIRAFFE HD: A High-Resolution 3D-aware Generative Model
3D-aware generative models have shown that the introduction of 3D information can lead to more controllable image generation. In particular, the current state-of-the-art model GIRAFFE can control each object's rotation, …
DisentanglementImage GenerationTranslationVocal Bursts Intensity Prediction