paper-with-me

Papers

Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning

2025-02-19 · Rui Zhao, Qirui Yuan, Jinyu Li, Haofeng Hu, Yun Li, Chengyuan Zheng, Fei Gao

End-to-end autonomous driving, which directly maps raw sensor inputs to low-level vehicle controls, is an important part of Embodied AI. Despite successes in applying Multimodal Large Language Models (MLLMs) for high-level traffic scene semantic understanding, it remains challenging to effectively translate these conceptual semantics understandings into low-level motion control commands and achieve generalization and consensus in cross-scene driving. We introduce Sce2DriveX, a human-like driving chain-of-thought (CoT) reasoning MLLM framework. Sce2DriveX utilizes multimodal joint learning from local scene videos and global BEV maps to deeply understand long-range spatiotemporal relationships and road topology, enhancing its comprehensive perception and reasoning capabilities in 3D dynamic/static scenes and achieving driving generalization across scenes. Building on this, it reconstructs the implicit cognitive chain inherent in human driving, covering scene understanding, meta-action reasoning, behavior interpretation analysis, motion planning and control, thereby further bridging the gap between autonomous driving and human thought processes. To elevate model performance, we have developed the first extensive Visual Question Answering (VQA) driving instruction dataset tailored for 3D spatial understanding and long-axis task reasoning. Extensive experiments demonstrate that Sce2DriveX achieves state-of-the-art performance from scene understanding to end-to-end driving, as well as robust generalization on the CARLA Bench2Drive benchmark.

📄 PDF Abstract BibTeX arXiv:2502.14917

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingBench2DriveMotion PlanningQuestion AnsweringScene UnderstandingVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Entropy Regularization 설명 없음
PPO Proximal Policy Optimization, or PPO, is a policy gradient method for reinforcement learning. The motivation was to have an algorithm with the data efficiency and reliable…
CARLA CARLA is an open-source simulator for autonomous driving research. CARLA has been developed from the ground up to support development, training, and validation of autonomous urban…

Similar Papers 제목 키워드 기반

DriveXQA: Cross-modal Visual Question Answering for Adverse Driving Scene Understanding

2026-03-11 · Mingzhe Tao, Ruiping Liu, Junwei Zheng, Yufan Chen 외 arxiv

Fusing sensors with complementary modalities is crucial for maintaining a stable and comprehensive understanding of abnormal driving scenes. However, Multimodal Large Language Models (MLLMs) are underexplored for leverag…

Visual Question AnsweringScene UnderstandingAutonomous VehiclesAutonomous Driving

DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving

2025-05-25 · Chen Shi, Shaoshuai Shi, Kehua Sheng, Bo Zhang 외

Data-driven learning has advanced autonomous driving, yet task-specific models struggle with out-of-distribution scenarios due to their narrow optimization objectives and reliance on costly annotated data. We present Dri…

Autonomous DrivingImage GenerationRepresentation LearningWorld Knowledge

Driving Scene Synthesis on Free-form Trajectories with Generative Prior

2024-12-02 · Zeyu Yang, Zijie Pan, Yuankun Yang, Xiatian Zhu 외

Driving scene synthesis along free-form trajectories is essential for driving simulations to enable closed-loop evaluation of end-to-end driving policies. While existing methods excel at novel view synthesis on recorded …

FormNovel View Synthesis

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

2026-06-28 · Sunqi Fan, Qingle Liu, Runqi Yin, Meng-Hao Guo 외 arxiv

Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Question Answering (VideoQA) benchmarks. Howe…

Video Question Answering

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

2025-12-14 · Zhe Liu, Runhui Huang, Rui Yang, Siming Yan 외 arxiv

Although multi-modal large language models (MLLMs) have shown strong capabilities across diverse domains, their application in generating fine-grained 3D perception and prediction outputs in autonomous driving remains un…

Autonomous DrivingPoint Clouds