MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving
In recent years, data-driven techniques have greatly advanced autonomous driving systems, but the need for rare and diverse training data remains a challenge, requiring significant investment in equipment and labor. World models, which predict and generate future environmental states, offer a promising solution by synthesizing annotated video data for training. However, existing methods struggle to generate long, consistent videos without accumulating errors, especially in dynamic scenes. To address this, we propose MiLA, a novel framework for generating high-fidelity, long-duration videos up to one minute. MiLA utilizes a Coarse-to-Re(fine) approach to both stabilize video generation and correct distortion of dynamic objects. Additionally, we introduce a Temporal Progressive Denoising Scheduler and Joint Denoising and Correcting Flow modules to improve the quality of generated videos. Extensive experiments on the nuScenes dataset show that MiLA achieves state-of-the-art performance in video generation quality. For more information, visit the project website: https://github.com/xiaomi-mlab/mila.github.io.
Code (1)
Tasks
Autonomous DrivingDenoisingVideo GenerationSimilar Papers 제목 키워드 기반
Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning
Text-to-multiview (T2MV) generation, which produces coherent multiview images from a single text prompt, remains computationally intensive, while accelerated T2MV methods using few-step diffusion models often sacrifice i…
Denoisingreinforcement-learningReinforcement LearningReinforcement Learning (RL)Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel Flow
Existing 3D editing methods rely on computationally intensive scene-by-scene iterative optimization and suffer from multi-view inconsistency. We propose an effective and feed-forward 3D editing framework based on the TRE…
Overview of Gaussian process based multi-fidelity techniques with variable relationship between fidelities
The design process of complex systems such as new configurations of aircraft or launch vehicles is usually decomposed in different phases which are characterized for instance by the depth of the analyses in terms of numb…
Gaussian ProcessesBeyond Semantic Similarity: A Two-Phase Non-Parametric Retrieval Workflow for Corporate Credit Underwriting
Corporate credit underwriting requires analysts to extract actionable evidence from long, heterogeneous financial documents spanning hundreds of pages and multiple languages. Standard Retrieval-Augmented Generation (RAG)…
Semantic SimilarityEnhanced Diagnostic Fidelity in Pathology Whole Slide Image Compression via Deep Learning
Accurate diagnosis of disease often depends on the exhaustive examination of Whole Slide Images (WSI) at microscopic resolution. Efficient handling of these data-intensive images requires lossy compression techniques. Th…
DiagnosticImage CompressionMS-SSIMSSIM+1