paper-with-me

Papers

RAG-Driver: Generalisable Driving Explanations with Retrieval-Augmented In-Context Learning in Multi-Modal Large Language Model

2024-02-16 · Jianhao Yuan, Shuyang Sun, Daniel Omeiza, Bo Zhao, Paul Newman, Lars Kunze, Matthew Gadd

We need to trust robots that use often opaque AI methods. They need to explain themselves to us, and we need to trust their explanation. In this regard, explainability plays a critical role in trustworthy autonomous decision-making to foster transparency and acceptance among end users, especially in complex autonomous driving. Recent advancements in Multi-Modal Large Language models (MLLMs) have shown promising potential in enhancing the explainability as a driving agent by producing control predictions along with natural language explanations. However, severe data scarcity due to expensive annotation costs and significant domain gaps between different datasets makes the development of a robust and generalisable system an extremely challenging task. Moreover, the prohibitively expensive training requirements of MLLM and the unsolved problem of catastrophic forgetting further limit their generalisability post-deployment. To address these challenges, we present RAG-Driver, a novel retrieval-augmented multi-modal large language model that leverages in-context learning for high-performance, explainable, and generalisable autonomous driving. By grounding in retrieved expert demonstration, we empirically validate that RAG-Driver achieves state-of-the-art performance in producing driving action explanations, justifications, and control signal prediction. More importantly, it exhibits exceptional zero-shot generalisation capabilities to unseen environments without further training endeavours.

📄 PDF Abstract BibTeX arXiv:2402.10828

Code (1)

YuanJianhao508/RAG-Driver pytorch

Tasks

Autonomous DrivingDecision MakingIn-Context LearningLanguage ModelingLanguage ModellingLarge Language ModelRAGRetrieval

Similar Papers 제목 키워드 기반

VLADriver-RAG: Retrieval-Augmented Vision-Language-Action Models for Autonomous Driving

2026-05-01 · Rui Zhao, Haofeng Hu, Zhenhai Gao, Jiaqiao Liu 외 arxiv

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving, yet their reliance on implicit parametric knowledge limits generalization in long-tail scenarios. While Retrieva…

Autonomous Driving

SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation

2025-07-29 · Hao Ye, Mengshi Qi, Zhaohong Liu, Liang Liu 외 arxiv

In this work, we study how vision-language models (VLMs) can be utilized to enhance the safety for the autonomous driving system, including perception, situational understanding, and path planning. However, existing rese…

Visual Question AnsweringInformation RetrievalAutonomous Driving

From Spoken Thoughts to Automated Driving Commentary: Predicting and Explaining Intelligent Vehicles' Actions

2022-04-19 · Daniel Omeiza, Sule Anjomshoae, Helena Webb, Marina Jirotka 외

In commentary driving, drivers verbalise their observations, assessments and intentions. By speaking out their thoughts, both learning and expert drivers are able to create a better understanding and awareness of their s…

counterfactual

Predicting Driver Fatigue in Automated Driving with Explainability

2021-03-03 · Feng Zhou, Areen Alsaid, Mike Blommer, Reates Curry 외

Research indicates that monotonous automated driving increases the incidence of fatigued driving. Although many prediction models based on advanced machine learning techniques were proposed to monitor driver fatigue, esp…

BIG-bench Machine Learning

Towards Safer and Understandable Driver Intention Prediction

2025-10-10 · Mukilan Karuppasamy, Shankar Gangisetty, Shyam Nandan Rai, Carlo Masone 외 arxiv

Autonomous driving (AD) systems are becoming increasingly capable of handling complex tasks, mainly due to recent advances in deep learning and AI. As interactions between autonomous systems and humans increase, the inte…

Action AnticipationAutonomous Driving