paper-with-me

홈 › Papers

Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps

2025-05-19 · Ziqi Wen, Jonathan Skaza, Shravan Murlidaran, William Y. Wang, Miguel P. Eckstein

Although models exist that predict human response times (RTs) in tasks such as target search and visual discrimination, the development of image-computable predictors for scene understanding time remains an open challenge. Recent advances in vision-language models (VLMs), which can generate scene descriptions for arbitrary images, combined with the availability of quantitative metrics for comparing linguistic descriptions, offer a new opportunity to model human scene understanding. We hypothesize that the primary bottleneck in human scene understanding and the driving source of variability in response times across scenes is the interaction between the foveated nature of the human visual system and the spatial distribution of task-relevant visual information within an image. Based on this assumption, we propose a novel image-computable model that integrates foveated vision with VLMs to produce a spatially resolved map of scene understanding as a function of fixation location (Foveated Scene Understanding Map, or F-SUM), along with an aggregate F-SUM score. This metric correlates with average (N=17) human RTs (r=0.47) and number of saccades (r=0.51) required to comprehend a scene (across 277 scenes). The F-SUM score also correlates with average (N=16) human description accuracy (r=-0.56) in time-limited presentations. These correlations significantly exceed those of standard image-based metrics such as clutter, visual complexity, and scene ambiguity based on language entropy. Together, our work introduces a new image-computable metric for predicting human response times in scene understanding and demonstrates the importance of foveated visual processing in shaping comprehension difficulty.

📄 PDF Abstract BibTeX arXiv:2505.12660

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Similar Papers 제목 키워드 기반

Foveation in the Era of Deep Learning

2023-12-03 · George Killick, Paul Henderson, Paul Siebert, Gerardo Aragon-Camarasa

In this paper, we tackle the challenge of actively attending to visual scenes using a foveated sensor. We introduce an end-to-end differentiable foveated active vision architecture that leverages a graph convolutional ne…

Deep LearningFoveationObject Recognition

Can Peripheral Representations Improve Clutter Metrics on Complex Scenes?

2016-08-14 · NeurIPS 2016 12 · Arturo Deza, Miguel P. Eckstein

Previous studies have proposed image-based clutter measures that correlate with human search times and/or eye movements. However, most models do not take into account the fact that the effects of clutter interact with th…

GazeProphet: Software-Only Gaze Prediction for VR Foveated Rendering

2025-08-19 · Farhaan Ebadulla, Chiraag Mudlapur, Gaurav BV arxiv

Foveated rendering significantly reduces computational demands in virtual reality applications by concentrating rendering quality where users focus their gaze. Current approaches require expensive hardware-based eye trac…

Medical Image Quality Metrics for Foveated Model Observers

2021-02-09 · Miguel A. Lago, Craig K. Abbey, Miguel P. Eckstein

A recently proposed model observer mimics the foveated nature of the human visual system by processing the entire image with varying spatial detail, executing eye movements and scrolling through slices. The model can pre…

model

Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding

2026-05-18 · Shravan Murlidaran, Ziqi Wen, Sana Shehabi, Miguel P. Eckstein arxiv

When humans view scenes without a specific task (free-viewing), they initially direct their eye movements toward the scene center and then fixate on people, text, objects being gazed at or grasped, and semantically meani…

Scene Understanding