paper-with-me

홈 › Papers

DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation

2024-09-09 · Wei Wu, Xi Guo, Weixuan Tang, Tingxuan Huang, Chiyu Wang, Dongyue Chen, Chenjing Ding

Recent advancements in generative models have provided promising solutions for synthesizing realistic driving videos, which are crucial for training autonomous driving perception models. However, existing approaches often struggle with multi-view video generation due to the challenges of integrating 3D information while maintaining spatial-temporal consistency and effectively learning from a unified model. We propose DriveScape, an end-to-end framework for multi-view, 3D condition-guided video generation, capable of producing 1024 x 576 high-resolution videos at 10Hz. Unlike other methods limited to 2Hz due to the 3D box annotation frame rate, DriveScape overcomes this with its ability to operate under sparse conditions. Our Bi-Directional Modulated Transformer (BiMot) ensures precise alignment of 3D structural information, maintaining spatial-temporal consistency. DriveScape excels in video generation performance, achieving state-of-the-art results on the nuScenes dataset with an FID score of 8.34 and an FVD score of 76.39. Our project homepage: https://metadrivescape.github.io/papers_project/drivescapev1/index.html

📄 PDF Abstract BibTeX arXiv:2409.05463

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingVideo Generation

Methods 이 논문이 사용한 방법론

Attention 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Position-Wise Feed-Forward Layer 설명 없음

Similar Papers 제목 키워드 기반

DriveScape: High-Resolution Driving Video Generation by Multi-View Feature Fusion

2025-01-01 · CVPR 2025 1 · Wei Wu, Xi Guo, Weixuan Tang, Tingxuan Huang 외

Recent advancements in generative models offer promising solutions for synthesizing realistic driving videos, aiding in training autonomous driving perception models. However, existing methods often struggle with hig…

Autonomous DrivingDenoisingVideo Generation

MyGo: Consistent and Controllable Multi-View Driving Video Generation with Camera Control

2024-09-10 · Yining Yao, Xi Guo, Chenjing Ding, Wei Wu

High-quality driving video generation is crucial for providing training data for autonomous driving models. However, current generative models rarely focus on enhancing camera motion control under multi-view tasks, which…

Autonomous DrivingVideo Generation

SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable Illumination

2026-09-09 · Athanasios Tragakis, Marco Aversa, Daniela Ivanova, Chaitanya Kaul 외 arxiv

SceneHI is a framework that lifts high-resolution, illumination-aware priors from 2D diffusion models to perform 3D texture synthesis. It is the first to demonstrate that high-resolution textures, previously limited to 2…

DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model

2026-06-29 · Daniyel Ayupov, Artur Markov-Tsoy arxiv

We present DreamForge-World 0.1 Preview, a preview foundational world model for real-time interactive world simulation. The system adapts the LongLive 1 autoregressive video stack, itself derived from Wan2.1-T2V-1.3B, wi…

GIRAFFE HD: A High-Resolution 3D-aware Generative Model

2022-03-28 · CVPR 2022 1 · Yang Xue, Yuheng Li, Krishna Kumar Singh, Yong Jae Lee

3D-aware generative models have shown that the introduction of 3D information can lead to more controllable image generation. In particular, the current state-of-the-art model GIRAFFE can control each object's rotation, …

DisentanglementImage GenerationTranslationVocal Bursts Intensity Prediction