paper-with-me

홈 › Papers

Disentangling top-down vs. bottom-up and low-level vs. high-level influences on eye movements over time

2018-05-17

Bottom-up and top-down, as well as low-level and high-level factors influence where we fixate when viewing natural scenes. However, the importance of each of these factors and how they interact remains a matter of debate. Here, we disentangle these factors by analysing their influence over time. For this purpose we develop a saliency model which is based on the internal representation of a recent early spatial vision model to measure the low-level bottom-up factor. To measure the influence of high-level bottom-up features, we use a recent DNN-based saliency model. To account for top-down influences, we evaluate the models on two large datasets with different tasks: first, a memorisation task and, second, a search task. Our results lend support to a separation of visual scene exploration into three phases: The first saccade, an initial guided exploration characterised by a gradual broadening of the fixation density, and an steady state which is reached after roughly 10 fixations. Saccade target selection during the initial exploration and in the steady state are related to similar areas of interest, which are better predicted when including high-level features. In the search dataset, fixation locations are determined predominantly by top-down processes. In contrast, the first fixation follows a different fixation density and contains a strong central fixation bias. Nonetheless, first fixations are guided strongly by image properties and as early as 200 ms after image onset, fixations are better predicted by high-level information. We conclude that any low-level bottom-up factors are mainly limited to the generation of the first saccade. All saccades are better explained when high-level features are considered, and later this high-level bottom-up control can be overruled by top-down influences.

📄 PDF Abstract BibTeX arXiv:1803.07352

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Disentangling neural mechanisms for perceptual grouping

2019-06-04 · ICLR 2020 1 · Junkyung Kim, Drew Linsley, Kalpit Thakkar, Thomas Serre

Forming perceptual groups and individuating objects in visual scenes is an essential step towards visual intelligence. This ability is thought to arise in the brain from computations implemented by bottom-up, horizontal,…

Object

An Iterative and Cooperative Top-Down and Bottom-Up Inference Network for Salient Object Detection

2019-06-01 · CVPR 2019 6 · Wenguan Wang, Jianbing Shen, Ming-Ming Cheng, Ling Shao

This paper presents a salient object detection method that integrates both top-down and bottom-up saliency inference in an iterative and cooperative manner. The top-down process is used for coarse-to-fine saliency estima…

object-detectionObject DetectionRGB Salient Object DetectionSaliency Prediction+1

Modulating Bottom-Up and Top-Down Visual Processing via Language-Conditional Filters

2020-03-28 · İlker Kesen, Ozan Arkan Can, Erkut Erdem, Aykut Erdem 외

How to best integrate linguistic and perceptual processing in multi-modal tasks that involve language and vision is an important open problem. In this work, we argue that the common practice of using language in a top-do…

ColorizationImage ColorizationReferring ExpressionReferring Expression Segmentation+1

Language Controls More Than Top-Down Attention: Modulating Bottom-Up Visual Processing with Referring Expressions

2021-01-01 · Ozan Arkan Can, Ilker Kesen, Deniz Yuret

How to best integrate linguistic and perceptual processing in multimodal tasks is an important open problem. In this work we argue that the common technique of using language to direct visual attention over high-level vi…

Referring Expression

MMLatch: Bottom-up Top-down Fusion for Multimodal Sentiment Analysis

2022-01-24 · Georgios Paraskevopoulos, Efthymios Georgiou, Alexandros Potamianos

Current deep learning approaches for multimodal fusion rely on bottom-up fusion of high and mid-level latent modality representations (late/mid fusion) or low level sensory inputs (early fusion). Models of human percepti…

Multimodal Sentiment AnalysisSentiment Analysis