paper-with-me

홈 › Papers

Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?

2016-06-17 · Abhishek Das, Harsh Agrawal, C. Lawrence Zitnick, Devi Parikh, Dhruv Batra

We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and test multiple game-inspired novel attention-annotation interfaces that require the subject to sharpen regions of a blurred image to answer a question. Thus, we introduce the VQA-HAT (Human ATtention) dataset. We evaluate attention maps generated by state-of-the-art VQA models against human attention both qualitatively (via visualizations) and quantitatively (via rank-order correlation). Overall, our experiments show that current attention models in VQA do not seem to be looking at the same regions as humans.

📄 PDF Abstract BibTeX arXiv:1606.05589

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?

2016-06-11 · EMNLP 2016 11 · Abhishek Das, Harsh Agrawal, C. Lawrence Zitnick, Devi Parikh 외

We conduct large-scale studies on `human attention' in Visual Question Answering (VQA) to understand where humans choose to look to answer questions about images. We design and test multiple game-inspired novel attention…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Differential Attention for Visual Question Answering

2018-04-01 · CVPR 2018 6 · Badri Patro, Vinay P. Namboodiri

In this paper we aim to answer questions based on images when provided with a dataset of question-answer pairs for a number of images during training. A number of methods have focused on solving this problem by using ima…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Towards Knowledge-Augmented Visual Question Answering

2020-12-01 · COLING 2020 8 · Maryam Ziaeefard, Freddy Lecue

Visual Question Answering (VQA) remains algorithmically challenging while it is effortless for humans. Humans combine visual observations with general and commonsense knowledge to answer questions about a given image. In…

General KnowledgeGraph AttentionQuestion AnsweringVisual Question Answering+1

On the Cognition of Visual Question Answering Models and Human Intelligence: A Comparative Study

2023-10-04 · Liben Chen, Long Chen, Tian Ellison-Chen, Zhuoyuan Xu

Visual Question Answering (VQA) is a challenging task that requires cross-modal understanding and reasoning of visual image and natural language question. To inspect the association of VQA models to human cognition, we d…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Optimizing Visual Question Answering Models for Driving: Bridging the Gap Between Human and Machine Attention Patterns

2024-06-13 · Kaavya Rekanar, Martin Hayes, Ganesh Sistu, Ciaran Eising

Visual Question Answering (VQA) models play a critical role in enhancing the perception capabilities of autonomous driving systems by allowing vehicles to analyze visual inputs alongside textual queries, fostering natura…

Autonomous DrivingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)