Interpretable Visual Question Answering by Visual Grounding from Attention Supervision Mining
A key aspect of VQA models that are interpretable is their ability to ground their answers to relevant regions in the image. Current approaches with this capability rely on supervised learning and human annotated groundings to train attention mechanisms inside the VQA architecture. Unfortunately, obtaining human annotations specific for visual grounding is difficult and expensive. In this work, we demonstrate that we can effectively train a VQA architecture with grounding supervision that can be automatically obtained from available region descriptions and object annotations. We also show that our model trained with this mined supervision generates visual groundings that achieve a higher correlation with respect to manually-annotated groundings, meanwhile achieving state-of-the-art VQA accuracy.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringVisual GroundingVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Equivariant and Invariant Grounding for Video Question Answering
Video Question Answering (VideoQA) is the task of answering the natural language questions about a video. Producing an answer requires understanding the interplay across visual scenes in video and linguistic semantics in…
Question AnsweringVideo Question AnsweringInterpretable Visual Question Answering via Reasoning Supervision
Transformer-based architectures have recently demonstrated remarkable performance in the Visual Question Answering (VQA) task. However, such models are likely to disregard crucial visual cues and often rely on multimodal…
Common Sense ReasoningQuestion AnsweringVisual GroundingVisual Question Answering+1Show Me What and Where has Changed? Question Answering and Grounding for Remote Sensing Change Detection
Remote sensing change detection aims to perceive changes occurring on the Earth's surface from remote sensing data in different periods, and feed these changes back to humans. However, most existing methods only focus on…
Change DetectionQuestion AnsweringVisual Question AnsweringVisDoT : Enhancing Visual Reasoning through Human-Like Interpretation Grounding and Decomposition of Thought
Large vision-language models (LVLMs) struggle to reliably detect visual primitives in charts and align them with semantic representations, which severely limits their performance on complex visual reasoning. This lack of…
Visual Question AnsweringVisual GroundingVisual ReasoningGrounding Answers for Visual Questions Asked by Visually Impaired People
Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQAGrounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual imp…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)