paper-with-me

홈 › Papers

InfinityDrive: Breaking Time Limits in Driving World Models

2024-12-02 · Xi Guo, Chenjing Ding, Haoxuan Dou, Xin Zhang, Weixuan Tang, Wei Wu

Autonomous driving systems struggle with complex scenarios due to limited access to diverse, extensive, and out-of-distribution driving data which are critical for safe navigation. World models offer a promising solution to this challenge; however, current driving world models are constrained by short time windows and limited scenario diversity. To bridge this gap, we introduce InfinityDrive, the first driving world model with exceptional generalization capabilities, delivering state-of-the-art performance in high fidelity, consistency, and diversity with minute-scale video generation. InfinityDrive introduces an efficient spatio-temporal co-modeling module paired with an extended temporal training strategy, enabling high-resolution (576$\times$1024) video generation with consistent spatial and temporal coherence. By incorporating memory injection and retention mechanisms alongside an adaptive memory curve loss to minimize cumulative errors, achieving consistent video generation lasting over 1500 frames (more than 2 minutes). Comprehensive experiments in multiple datasets validate InfinityDrive's ability to generate complex and varied scenarios, highlighting its potential as a next-generation driving world model built for the evolving demands of autonomous driving. Our project homepage: https://metadrivescape.github.io/papers_project/InfinityDrive/page.html

📄 PDF Abstract BibTeX arXiv:2412.01522

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingDiversityVideo Generation

Similar Papers 제목 키워드 기반

Symbolic Imitation Learning: From Black-Box to Explainable Driving Policies

2023-09-27 · Iman Sharifi, Saber Fallah

Current methods of imitation learning (IL), primarily based on deep neural networks, offer efficient means for obtaining driving policies from real-world data but suffer from significant limitations in interpretability a…

Autonomous DrivingImitation LearningInductive logic programming

Pedestrian and Ego-Vehicle Trajectory Prediction From Monocular Camera

2021-06-19 · CVPR 2021 1 · Lukas Neumann, Andrea Vedaldi

Predicting future pedestrian trajectory is a crucial component of autonomous driving systems, as recognizing critical situations based only on current pedestrian position may come too late for any meaningful correcti…

Autonomous DrivingPedestrian Trajectory PredictionPose PredictionPosition+1

Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking

2025-07-06 · Aldan Creo, Raul Castro Fernandez, Manuel Cebrian arxiv

As large language models (LLMs) become increasingly deployed, understanding the complexity and evolution of jailbreaking strategies is critical for AI safety. We present a mass-scale empirical analysis of jailbreak compl…

LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model

2025-06-02 · Xiaodong Wang, Zhirong Wu, Peixi Peng

Driving world models are used to simulate futures by video generation based on the condition of the current state and actions. However, current models often suffer serious error accumulations when predicting the long-ter…

Video Generation

Preventing Robotic Jailbreaking via Multimodal Domain Adaptation

2025-09-27 · Francesco Marchiori, Rohan Sinha, Christopher Agia, Alexander Robey 외 arxiv

Large Language Models (LLMs) and Vision-Language Models (VLMs) are increasingly deployed in robotic environments but remain vulnerable to jailbreaking attacks that bypass safety mechanisms and drive unsafe or physically …

Autonomous DrivingDomain Adaptation