paper-with-me

홈 › Papers

VQA-MHUG: A Gaze Dataset to Study Multimodal Neural Attention in Visual Question Answering

2021-09-27 · CoNLL (EMNLP) 2021 11 · Ekta Sood, Fabian Kögel, Florian Strohm, Prajit Dhar, Andreas Bulling

We present VQA-MHUG - a novel 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker. We use our dataset to analyze the similarity between human and neural attentive strategies learned by five state-of-the-art VQA models: Modular Co-Attention Network (MCAN) with either grid or region features, Pythia, Bilinear Attention Network (BAN), and the Multimodal Factorized Bilinear Pooling Network (MFB). While prior work has focused on studying the image modality, our analyses show - for the first time - that for all models, higher correlation with human attention on text is a significant predictor of VQA performance. This finding points at a potential for improving VQA performance and, at the same time, calls for further research on neural text attention mechanisms and their integration into architectures for vision and language tasks, including but potentially also beyond VQA.

📄 PDF Abstract BibTeX arXiv:2109.13116

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Look Hear: Gaze Prediction for Speech-directed Human Attention

2024-07-28 · Sounak Mondal, Seoyoung Ahn, Zhibo Yang, Niranjan Balasubramanian 외

For computer systems to effectively interact with humans using spoken language, they need to understand how the words being generated affect the users' moment-by-moment attention. Our study focuses on the incremental pre…

DecoderGaze PredictionReferring ExpressionScanpath prediction

The AICO Multimodal Corpus -- Data Collection and Preliminary Analyses

2020-05-01 · LREC 2020 5 · Kristiina Jokinen

This paper describes data collection and the first explorative research on the AICO Multimodal Corpus. The corpus contains eye-gaze, Kinect, and video recordings of human-robot and human-human interactions, and was colle…

Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models

2026-05-19 · Hengfei Wang, Anshul Gupta, Pierre Vuillecard, Jean-Marc Odobez arxiv

Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs could greatly benefit the analysis of human gaze and attention, a c…

Zero-shot GeneralizationRelational Reasoning

CSGaze: Context-aware Social Gaze Prediction

2025-11-08 · Surbhi Madan, Shreya Ghosh, Ramanathan Subramanian, Abhinav Dhall 외 arxiv

A person's gaze offers valuable insights into their focus of attention, level of social engagement, and confidence. In this work, we investigate how contextual cues combined with visual scene and facial information can b…

A Modular Multimodal Architecture for Gaze Target Prediction: Application to Privacy-Sensitive Settings

2023-07-11 · IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops 2022 6 · Anshul Gupta, Samy Tafasca, Jean-Marc Odobez

Predicting where a person is looking is a complex task, requiring to understand not only the person's gaze and scene content, but also the 3D scene structure and the person's situation (are they manipulating? interacting…