paper-with-me

홈 › Papers

DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving

2024-08-29 · Yongjie Fu, Anmol Jain, Xuan Di, Xu Chen, Zhaobin Mo

The advancement of autonomous driving technologies necessitates increasingly sophisticated methods for understanding and predicting real-world scenarios. Vision language models (VLMs) are emerging as revolutionary tools with significant potential to influence autonomous driving. In this paper, we propose the DriveGenVLM framework to generate driving videos and use VLMs to understand them. To achieve this, we employ a video generation framework grounded in denoising diffusion probabilistic models (DDPM) aimed at predicting real-world video sequences. We then explore the adequacy of our generated videos for use in VLMs by employing a pre-trained model known as Efficient In-context Learning on Egocentric Videos (EILEV). The diffusion model is trained with the Waymo open dataset and evaluated using the Fr\'echet Video Distance (FVD) score to ensure the quality and realism of the generated videos. Corresponding narrations are provided by EILEV for these generated videos, which may be beneficial in the autonomous driving domain. These narrations can enhance traffic scene understanding, aid in navigation, and improve planning capabilities. The integration of video generation with VLMs in the DriveGenVLM framework represents a significant step forward in leveraging advanced AI models to address complex challenges in autonomous driving.

📄 PDF Abstract BibTeX arXiv:2408.16647

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingDenoisingIn-Context LearningLanguage ModelingLanguage ModellingScene UnderstandingVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation

2025-10-01 · Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li 외 arxiv

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Visi…

Trajectory ForecastingTrajectory PredictionVideo Generation

Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation

2026-02-05 · Hai Zhang, Siqi Liang, Li Chen, Yuxian Li 외 arxiv

Why must vision-language navigation be bound to detailed and verbose language instructions? While such details ease decision-making, they fundamentally contradict the goal for navigation in the real-world. Ideally, agent…

Vision-Language NavigationVideo Generation

VideoAgent: Self-Improving Video Generation

2024-10-14 · Achint Soni, Sreyas Venkataraman, Abhranil Chandra, Sebastian Fischmeister 외

Video generation has been used to generate visual plans for controlling robotic systems. Given an image observation and a language instruction, previous work has generated video plans which are then converted to robot co…

HallucinationVideo Generation

WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling

2025-12-08 · Shaoheng Fang, Hanwen Jiang, Yunpeng Bai, Niloy J. Mitra 외 arxiv

Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally consistent. WorldReel jointly produces RGB…

Video Generation

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

2024-12-10 · CVPR 2025 1 · Xi Chen, Zhifei Zhang, He Zhang, Yuqian Zhou 외

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs an…

Image GenerationVideo Generation