Modeling the Contribution of Central Versus Peripheral Vision in Scene, Object, and Face Recognition
It is commonly believed that the central visual field is important for recognizing objects and faces, and the peripheral region is useful for scene recognition. However, the relative importance of central versus peripheral information for object, scene, and face recognition is unclear. In a behavioral study, Larson and Loschky (2009) investigated this question by measuring the scene recognition accuracy as a function of visual angle, and demonstrated that peripheral vision was indeed more useful in recognizing scenes than central vision. In this work, we modeled and replicated the result of Larson and Loschky (2009), using deep convolutional neural networks. Having fit the data for scenes, we used the model to predict future data for large-scale scene recognition as well as for objects and faces. Our results suggest that the relative order of importance of using central visual field information is face recognition>object recognition>scene recognition, and vice-versa for peripheral information.
Code (0)
등록된 구현이 없습니다.
Tasks
Face RecognitionObject RecognitionScene RecognitionSimilar Papers 제목 키워드 기반
Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality
The performance of modern language models depends critically on pretraining data composition. Yet existing data selection methods rely on auxiliary classifiers for document scoring or mixture optimization, adding computa…
How spatial frequencies and color drive object search in real-world scenes: A new eye-movement corpus
When studying how people search for objects in scenes, the inhomogeneity of the visual field is often ignored. Due to physiological limitations peripheral vision is blurred and mainly uses coarse-grained information (i.e…
ObjectObject LocalizationCVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
We present a central-peripheral vision-inspired framework (CVP), a simple yet effective multimodal model for spatial reasoning that draws inspiration from the two types of human visual fields -- central vision and periph…
Scene UnderstandingSpatial ReasoningPoint CloudsCross-Temporal Attention Fusion (CTAF) for Multimodal Physiological Signals in Self-Supervised Learning
We study multimodal affect modeling when EEG and peripheral physiology are asynchronous, which most fusion methods ignore or handle with costly warping. We propose Cross-Temporal Attention Fusion (CTAF), a self-supervise…
Self-Supervised LearningEstimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games
We study how different visual information sources contribute to human decision making in dynamic visual environments. Using Atari-HEAD, a large-scale Atari gameplay dataset with synchronized eye-tracking, we introduce a …
Decision MakingAtari Games