paper-with-me

홈 › Papers

Eye Gaze and Self-attention: How Humans and Transformers Attend Words in Sentences

2022-05-01 · CMCL (ACL) 2022 5 · Joshua Bensemann, Alex Peng, Diana Prado, Yang Chen, Neset Tan, Paul Michael Corballis, Patricia Riddle, Michael Witbrock

Attention describes cognitive processes that are important to many human phenomena including reading. The term is also used to describe the way in which transformer neural networks perform natural language processing. While attention appears to be very different under these two contexts, this paper presents an analysis of the correlations between transformer attention and overt human attention during reading tasks. An extensive analysis of human eye tracking datasets showed that the dwell times of human eye movements were strongly correlated with the attention patterns occurring in the early layers of pre-trained transformers such as BERT. Additionally, the strength of a correlation was not related to the number of parameters within a transformer. This suggests that something about the transformers’ architecture determined how closely the two measures were correlated.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Eye Gaze and Self-attention: How Humans and Transformers Attend Words in Sentences

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Attention mechanisms are used to describe human reading processes and natural language processing by transformer neural networks. On the surface, attention appears to be very different under these two contexts. However, …

Gaze Perception in Humans and CNN-Based Model

2021-04-17 · Nicole X. Han, William Yang Wang, Miguel P. Eckstein

Making accurate inferences about other individuals' locus of attention is essential for human social interactions and will be important for AI to effectively interact with humans. In this study, we compare how a CNN (con…

model

Do humans and Convolutional Neural Networks attend to similar areas during scene classification: Effects of task and image type

2023-07-25 · Romy Müller, Marcel Dürschmidt, Julian Ullrich, Carsten Knoll 외

Deep Learning models like Convolutional Neural Networks (CNN) are powerful image classifiers, but what factors determine whether they attend to similar image areas as humans do? While previous studies have focused on tec…

Explainable artificial intelligenceScene Classification

ViTGaze: Gaze Following with Interaction Features in Vision Transformers

2024-03-19 · Yuehao Song, Xinggang Wang, Jingfeng Yao, Wenyu Liu 외

Gaze following aims to interpret human-scene interactions by predicting the person's focal point of gaze. Prevailing approaches often adopt a two-stage framework, whereby multi-modality information is extracted in the in…

Gaze Target Estimation

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models

2026-05-13 · Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo 외 arxiv

When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their intended description. In contrast, current multimodal large language mode…