paper-with-me

홈 › Papers

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior

2025-07-24 · Junda Wu, Jessica Echterhoff, Kyungtae Han, Amr Abdelraouf, Rohit Gupta, Julian McAuley arxiv

Understanding a driver's behavior and intentions is important for potential risk assessment and early accident prevention. Safety and driver assistance systems can be tailored to individual drivers' behavior, significantly enhancing their effectiveness. However, existing datasets are limited in describing and explaining general vehicle movements based on external visual evidence. This paper introduces a benchmark, PDB-Eval, for a detailed understanding of Personalized Driver Behavior, and aligning Large Multimodal Models (MLLMs) with driving comprehension and reasoning. Our benchmark consists of two main components, PDB-X and PDB-QA. PDB-X can evaluate MLLMs' understanding of temporal driving scenes. Our dataset is designed to find valid visual evidence from the external view to explain the driver's behavior from the internal view. To align MLLMs' reasoning abilities with driving tasks, we propose PDB-QA as a visual explanation question-answering task for MLLM instruction fine-tuning. As a generic learning task for generative models like MLLMs, PDB-QA can bridge the domain gap without harming MLLMs' generalizability. Our evaluation indicates that fine-tuning MLLMs on fine-grained descriptions and explanations can effectively bridge the gap between MLLMs and the driving domain, which improves zero-shot performance on question-answering tasks by up to 73.2%. We further evaluate the MLLMs fine-tuned on PDB-X in Brain4Cars' intention prediction and AIDE's recognition tasks. We observe up to 12.5% performance improvements on the turn intention prediction task in Brain4Cars, and consistent performance improvements up to 11.0% on all tasks in AIDE.

📄 PDF Abstract BibTeX arXiv:2507.18447

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models

2026-06-15 · Yupei Li, Qiyang Sun, Xiaoliang Wu, Chenxi Wang 외 arxiv

Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fall into two categories. Traditional explainable AI (XAI), such as gradient-based a…

Explanation GenerationDeepFake Detection

An Experimental Study on Generating Plausible Textual Explanations for Video Summarization

2025-09-30 · Thomas Eleftheriadis, Evlampios Apostolidis, Vasileios Mezaris arxiv

In this paper, we present our experimental study on generating plausible textual explanations for the outcomes of video summarization. For the needs of this study, we extend an existing framework for multigranular explan…

Video Summarization

LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation

2026-05-08 · Jun Wang, Fengpeng Li, Hang Dong, Tianjin Huang 외 arxiv

Remote sensing lithology interpretation is fundamental to geological surveys, mineral exploration, and regional geological mapping. Unlike general land-cover recognition, lithology interpretation is a knowledge-intensive…

Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier

2025-10-27 · Hyeongseop Rha, Jeong Hun Yeo, Yeonju Kim, Yong Man Ro arxiv

The recent advancement of Multimodal Large Language Models (MLLMs) is transforming human-computer interaction (HCI) from surface-level exchanges into more nuanced and emotionally intelligent communication. To realize thi…

Multimodal Emotion Recognition

Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering

2025-09-28 · Eduard Barbu, Adrian Marius Dumitran arxiv

Ensuring that both new and experienced drivers master current traffic rules is critical to road safety. This paper evaluates Large Language Models (LLMs) on Romanian driving-law QA with explanation generation. We release…

Explanation GenerationQuestion Answering