paper-with-me

홈 › Papers

I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automotive Contexts

2026-06-12 · Shaun Feakins, Ibrahim Habli, Kim Littler, Robert Palin arxiv

This paper appraises recent frameworks within AI development to integrate LLMs into control tasks in automotive contexts from the perspective of safety assurance. This work has built upon the rapid integration of LLMs across automotive settings. However, we find that at present, these frameworks face significant challenges, limiting their efficacy in real-time safety-critical contexts. Firstly, we consider conceptual challenges, including the fact that deployers are faced with a dual challenge, wherein they must assure a model which has been developed upstream, i.e. as general-purpose tools by the large AI labs, in a downstream context, i.e. into specific vehicle architectures. Secondly, we consider concrete challenges from across existing standards. We show that there are currently both fundamental engineering constraints covered in ISO21448, such as latency, and novel LLM-specific issues, such as alignment-related issues covered in ISO/PAS8800. We ground both examples in a concrete introductory, experimental case study exploring an existing open-source repository, Talk2Drive. We present a safety argument in order to make explicit the limitations of existing solutions. Nonetheless, given that the use of LLMs in automotive contexts is being explored at a technical level and operationalised, we propose potential assurance mechanisms for LLM-related hazardous events going forward.

📄 PDF Abstract BibTeX arXiv:2606.14327

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)

2024-08-28 · Aman Priyanshu, Supriti Vijay

This paper introduces FRACTURED-SORRY-Bench, a framework for evaluating the safety of Large Language Models (LLMs) against multi-turn conversational attacks. Building upon the SORRY-Bench dataset, we propose a simple yet…

SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors

2024-06-20 · Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang 외

Evaluating aligned large language models' (LLMs) ability to recognize and reject unsafe user requests is crucial for safe, policy-compliant deployments. Existing evaluation efforts, however, face three limitations that w…

Language ModelingLanguage ModellingLarge Language Model

I'm sorry Dave, I'm afraid I can't do that, Deep Q-learning from forbidden action

2019-10-04 · Mathieu Seurin, Philippe Preux, Olivier Pietquin

The use of Reinforcement Learning (RL) is still restricted to simulation or to enhance human-operated systems through recommendations. Real-world environments (e.g. industrial robots or power grids) are generally designe…

Industrial RobotsQ-LearningReinforcement LearningReinforcement Learning (RL)+1

Gradient-Controlled Decoding: A Safety Guardrail for LLMs with Dual-Anchor Steering

2026-04-06 · Purva Chiniya, Kevin Scaria, Sagar Chaturvedi arxiv

Large language models (LLMs) remain susceptible to jailbreak and direct prompt-injection attacks, yet the strongest defensive filters frequently over-refuse benign queries and degrade user experience. Previous work on ja…

I'm Sorry Dave, I'm Afraid I Can't Return That: On YouTube Search API Use in Research

2025-06-04 · Alexandros Efstratiou

YouTube is among the most widely-used platforms worldwide, and has seen a lot of recent academic attention. Despite its popularity and the number of studies conducted on it, much less is understood about the way in which…