paper-with-me

홈 › Papers

Comparing Human Gaze and Vision-Language Model Attention in Safety-Relevant Environments

2026-06-13 · Marta Vallejo, Siwen Wang arxiv

Human visual attention plays an important role in how people perceive and respond to environments containing potential risks. This study investigates whether large vision-language models can identify the same regions of a scene that attract human attention in safety-relevant environments. Eye-tracking data were collected from ten participants viewing 33 scene images representing environments with varying levels of potential risk using Pupil Invisible wearable glasses. Gaze coordinates were mapped onto stimulus images to generate population-averaged human gaze heatmaps. In parallel, GPT-4o was prompted through the OpenAI Vision Application Programming Interface (API) to generate spatial predictions of visual attention, which were converted into saliency maps for comparison with human gaze patterns. Spatial alignment between human gaze heatmaps and model-generated saliency maps was evaluated using four complementary metrics: Pearson correlation (r = 0.515 +- 0.117), Normalised Scanpath Saliency (NSS = 0.988 +- 0.323), Kullback-Leibler divergence (KL = 1.766 +- 0.844), and Area Under the Receiver Operating Characteristic Curve using the Judd formulation (AUC-Judd = 0.806 +- 0.076). A cross-model comparison with Gemini Pro, Gemini Flash, and Claude showed that all models exceeded the AUC-Judd chance baseline of 0.5 and achieved positive NSS scores. Gemini Pro demonstrated the strongest spatial localisation according to three of the four metrics, whereas GPT-4o produced the closest distributional match to human attention as measured by KL divergence. These findings suggest that large vision-language models can identify regions that broadly correspond to where humans direct visual attention in safety-relevant scenes without requiring eye-tracking training data. The results highlight the potential of vision-language models as a scalable tool for approximating human attentional patterns.

📄 PDF Abstract BibTeX arXiv:2606.15202

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Gaze-Informed Vision Transformers: Predicting Driving Decisions Under Uncertainty

2023-08-26 · Sharath Koorathota, Nikolas Papadopoulos, Jia Li Ma, Shruti Kumar 외

Vision Transformers (ViT) have advanced computer vision, yet their efficacy in complex tasks like driving remains less explored. This study enhances ViT by integrating human eye gaze, captured via eye-tracking, to increa…

Autonomous Driving

Voila-A: Aligning Vision-Language Models with User's Gaze Attention

2023-12-22 · Kun Yan, Lei Ji, Zeyu Wang, Yuntao Wang 외

In recent years, the integration of vision and language understanding has led to significant advancements in artificial intelligence, particularly through Vision-Language Models (VLMs). However, existing VLMs face challe…

Interpretable Modeling of Driver Attention Shifts with a Vision-Language Model

2025-08-07 · Kaiser Hamid, Khandakar Ashrafi Akbar, Peihang Li, Nade Liang arxiv

Driver gaze is commonly modeled as a spatial heatmap, but heatmaps alone are difficult for humans to interpret because they do not explain which road object or region is being monitored or why an attention shift may matt…

Entity Alignment

Visual Attention Prediction Improves Performance of Autonomous Drone Racing Agents

2022-01-07 · Christian Pfeiffer, Simon Wengeler, Antonio Loquercio, Davide Scaramuzza

Humans race drones faster than neural networks trained for end-to-end autonomous flight. This may be related to the ability of human pilots to select task-relevant visual information effectively. This work investigates w…

Decision MakingImitation LearningPrediction

Generating Image Descriptions via Sequential Cross-Modal Alignment Guided by Human Gaze

2020-11-09 · EMNLP 2020 11 · Ece Takmaz, Sandro Pezzelle, Lisa Beinborn, Raquel Fernández

When speakers describe an image, they tend to look at objects before mentioning them. In this paper, we investigate such sequential cross-modal alignment by modelling the image description generation process computationa…

cross-modal alignmentImage CaptioningImage Description