paper-with-me

홈 › Papers

DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving

2025-05-27 · Muxi Diao, Lele Yang, Hongbo Yin, Zhexu Wang, Yejie Wang, Daxin Tian, Kongming Liang, Zhanyu Ma

Autonomous driving requires real-time, robust reasoning across perception, prediction, planning, and behavior. However, conventional end-to-end models fail to generalize in complex scenarios due to the lack of structured reasoning. Recent vision-language models (VLMs) have been applied to driving tasks, but they typically rely on isolated modules and static supervision, limiting their ability to support multi-stage decision-making. We present AutoDriveRL, a unified training framework that formulates autonomous driving as a structured reasoning process over four core tasks. Each task is independently modeled as a vision-language question-answering problem and optimized using task-specific reward models, enabling fine-grained reinforcement signals at different reasoning stages. Within this framework, we train DriveRX, a cross-task reasoning VLM designed for real-time decision-making. DriveRX achieves strong performance on a public benchmark, outperforming GPT-4o in behavior reasoning and demonstrating robustness under complex or corrupted driving conditions. Our analysis further highlights the impact of vision encoder design and reward-guided reasoning compression. We will release the AutoDriveRL framework and the DriveRX model to support future research.

📄 PDF Abstract BibTeX arXiv:2505.20665

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous DrivingDecision MakingQuestion Answering

Similar Papers 제목 키워드 기반

R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

2025-03-13 · Yi Yang, Xiaoxuan He, Hongkun Pan, Xiyan Jiang 외

Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge.…

Multimodal Reasoning

Learning Multi-View Spatial Reasoning from Cross-View Relations

2026-03-30 · Suchae Jeong, Jaehwi Song, Haeone Lee, Hanna Kim 외 arxiv

Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems to understand 3D environments and manip…

Spatial Reasoning

Code Execution as Grounded Supervision for LLM Reasoning

2025-06-12 · Dongwon Jung, Wenxuan Zhou, Muhao Chen

Training large language models (LLMs) with chain-of-thought (CoT) supervision has proven effective for enhancing their reasoning abilities. However, obtaining reliable and accurate reasoning supervision remains a signifi…

Dataset Generation

Mind to Hand: Purposeful Robotic Control via Embodied Reasoning

2025-12-09 · Peijun Tang, Shangjin Xie, Binyan Sun, Baifu Huang 외 arxiv

Humans act with context and intention, with reasoning playing a central role. While internet-scale data has enabled broad reasoning capabilities in AI systems, grounding these abilities in physical action remains a major…

Reinforcement LearningTrajectory Prediction

Cross-Modality Relevance for Reasoning on Language and Vision

2020-05-12 · ACL 2020 6 · Chen Zheng, Quan Guo, Parisa Kordjamshidi

This work deals with the challenge of learning and reasoning over language and vision data for the related downstream tasks such as visual question answering (VQA) and natural language for visual reasoning (NLVR). We des…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning