paper-with-me

홈 › Papers

Interpretable Modeling of Driver Attention Shifts with a Vision-Language Model

2025-08-07 · Kaiser Hamid, Khandakar Ashrafi Akbar, Peihang Li, Nade Liang arxiv

Driver gaze is commonly modeled as a spatial heatmap, but heatmaps alone are difficult for humans to interpret because they do not explain which road object or region is being monitored or why an attention shift may matter. This study examines whether minimal human-grounded supervision can steer a vision--language model toward interpretable descriptions of driver attention shifts. Using selected high-change gaze moments from the Berkeley DeepDrive-Attention dataset, we compare zero-shot, one-shot, and LoRA fine-tuned VLM conditions against human-refined reference descriptions and expert ratings. Results show that fine-tuning with 80 expert-refined attention examples improves ROUGE-L, METEOR, Entity Alignment F1, and Human Alignment Score relative to unsteered VLM outputs. The findings suggest that language-based descriptions can complement gaze heatmaps by making driver attention more accessible for human-factors analysis, driver-monitoring review, and situation-awareness support.

📄 PDF Abstract BibTeX arXiv:2508.05852

Code (0)

등록된 구현이 없습니다.

Tasks

Entity Alignment

Similar Papers 제목 키워드 기반

FSDAM: Few-Shot Driving Attention Modeling via Vision-Language Coupling

2025-11-16 · Kaiser Hamid, Can Cui, Khandakar Ashrafi Akbar, Ziran Wang 외 arxiv

Understanding not only where drivers look but also why their attention shifts is essential for interpretable human-AI collaboration in autonomous driving. Driver attention is not purely perceptual but semantically struct…

Zero-shot GeneralizationExplanation GenerationAutonomous Driving

A Semantic Decoupling-Based Two-Stage Rainy-Day Attack for Revealing Weather Robustness Deficiencies in Vision-Language Models

2026-01-19 · Chengyin Hu, Xiang Chen, Zhe Jia, Weiwen Shi 외 arxiv

Vision-Language Models (VLMs) are trained on image-text pairs collected under canonical visual conditions and achieve strong performance on multimodal tasks. However, their robustness to real-world weather conditions, an…

DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning

2026-03-30 · Weimin Liu, Qingkun Li, Jiyuan Qiu, Wenjun Wang 외 arxiv

Drivers' visual attention provides critical cues for anticipating latent hazards and directly shapes decision-making and control maneuvers, where its absence can compromise traffic safety. To emulate drivers' perception …

Scene Understanding

From Scene to Object: Text-Guided Dual-Gaze Prediction

2026-04-22 · Zehong Ke, Yanbo Jiang, Jinhao Li, Zhiyuan Liu 외 arxiv

Interpretable driver attention prediction is crucial for human-like autonomous driving. However, existing datasets provide only scene-level global gaze rather than fine-grained object-level annotations, inherently failin…

Autonomous Driving

Weather-Dependent Variations in Driver Gaze Behavior: A Case Study in Rainy Conditions

2025-08-31 · Ghazal Farhani, Taufiq Rahman, Dominique Charlebois arxiv

Rainy weather significantly increases the risk of road accidents due to reduced visibility and vehicle traction. Understanding how experienced drivers adapt their visual perception through gaze behavior under such condit…