paper-with-me

홈 › Papers

CARE Drive A Framework for Evaluating Reason-Responsiveness of Vision Language Models in Automated Driving

2026-02-17 · Lucas Elbert Suryana, Farah Bierenga, Sanne van Buuren, Pepijn Kooij, Elsefien Tulleners, Federico Scari, Simeon Calvert, Bart van Arem, Arkady Zgonnikov arxiv

Foundation models, including vision language models, are increasingly used in automated driving to interpret scenes, recommend actions, and generate natural language explanations. However, existing evaluation methods primarily assess outcome based performance, such as safety and trajectory accuracy, without determining whether model decisions reflect human relevant considerations. As a result, it remains unclear whether explanations produced by such models correspond to genuine reason responsive decision making or merely post hoc rationalizations. This limitation is especially significant in safety critical domains because it can create false confidence. To address this gap, we propose CARE Drive, Context Aware Reasons Evaluation for Driving, a model agnostic framework for evaluating reason responsiveness in vision language models applied to automated driving. CARE Drive compares baseline and reason augmented model decisions under controlled contextual variation to assess whether human reasons causally influence decision behavior. The framework employs a two stage evaluation process. Prompt calibration ensures stable outputs. Systematic contextual perturbation then measures decision sensitivity to human reasons such as safety margins, social pressure, and efficiency constraints. We demonstrate CARE Drive in a cyclist overtaking scenario involving competing normative considerations. Results show that explicit human reasons significantly influence model decisions, improving alignment with expert recommended behavior. However, responsiveness varies across contextual factors, indicating uneven sensitivity to different types of reasons. These findings provide empirical evidence that reason responsiveness in foundation models can be systematically evaluated without modifying model parameters.

📄 PDF Abstract BibTeX arXiv:2602.15645

Code (0)

등록된 구현이 없습니다.

Tasks

Decision Making

Similar Papers 제목 키워드 기반

When Clients Stop Following: A Cognitive Conceptualization Diagram-driven Framework for Strategic Counseling

2026-06-03 · Yihao Qin, Junyi Zhao, Changsheng Ma, Yongfeng Tao 외 arxiv

Large Language Models (LLMs) show promise in psychological counseling, yet existing benchmarks rely heavily on highly cooperative simulated clients. We observe a critical counselor-following phenomenon: these clients oft…

Reinforcement LearningResponse Generation

Cultural Prompting Improves the Empathy and Cultural Responsiveness of GPT-Generated Therapy Responses

2025-10-19 · Serena Jinchen Xie, Shumenghui Zhai, Yanjing Liang, Jingyi Li 외 arxiv

Large Language Model (LLM)-based conversational agents offer promising solutions for mental health support, but lack cultural responsiveness for diverse populations. This study evaluated the effectiveness of cultural pro…

MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables

2025-09-15 · Matteo Marcuzzo, Alessandro Zangari, Andrea Albarelli, Jose Camacho-Collados 외 arxiv

As LLMs excel on standard reading comprehension benchmarks, attention is shifting toward evaluating their capacity for complex abstract reasoning and inference. Literature-based benchmarks, with their rich narrative and …

Reading ComprehensionQuestion Answering

XR-CareerAssist: An Immersive Platform for Personalised Career Guidance Leveraging Extended Reality and Multimodal AI

2026-04-08 · N. D. Tantaroudas, A. J. McCracken, I. Karachalios, E. Papatheou 외 arxiv

Conventional career guidance platforms rely on static, text-driven interfaces that struggle to engage users or deliver personalised, evidence-based insights. Although Computer-Assisted Career Guidance Systems have evolve…

Machine TranslationSpeech Recognition

Designing Explainable AI for Healthcare Reviews: Guidance on Adoption and Trust

2026-02-11 · Eman Alamoudi, Ellis Solaiman arxiv

Patients increasingly rely on online reviews when choosing healthcare providers, yet the sheer volume of these reviews can hinder effective decision-making. This paper summarises a mixed-methods study aimed at evaluating…