Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering
Accurately answering a question about a given image requires combining
observations with general knowledge. While this is effortless for humans,
reasoning with general knowledge remains an algorithmic challenge. To advance
research in this direction a novel fact-based' visual question answering
(FVQA) task has been introduced recently along with a large set of curated
facts which link two entities, i.e., two possible answers, via a relation.
Given a question-image pair, deep network techniques have been employed to
successively reduce the large set of facts until one of the two entities of the
final remaining fact is predicted as the answer. We observe that a successive
process which considers one fact at a time to form a local decision is
sub-optimal. Instead, we develop an entity graph and use a graph convolutional
network to reason' about the correct answer by jointly considering all
entities. We show on the challenging FVQA dataset that this leads to an
improvement in accuracy of around 7% compared to the state of the art.
Code (0)
등록된 구현이 없습니다.
Tasks
Factual Visual Question AnsweringGeneral KnowledgeQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Iterative Visual Reasoning Beyond Convolutions
We present a novel framework for iterative visual reasoning. Our framework goes beyond current recognition systems that lack the capability to reason beyond stack of convolutions. The framework consists of two core modul…
Visual ReasoningMucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering
Fact-based Visual Question Answering (FVQA) requires external knowledge beyond visible content to answer questions about an image, which is challenging but indispensable to achieve general VQA. One limitation of existing…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
Chart question-answering (QA) benchmarks aim to pose questions that require visual reasoning to correctly answer, but models can often reach solutions through shortcuts or prior familiarity with a chart based on their ow…
Visual ReasoningBeyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models
Visual Language Models (VLMs) are powerful generative tools but often produce factually inaccurate outputs due to a lack of robust reasoning capabilities. While extensive research has been conducted on integrating extern…
Knowledge GraphsCross-modal Knowledge Reasoning for Knowledge-based Visual Question Answering
Knowledge-based Visual Question Answering (KVQA) requires external knowledge beyond the visible content to answer questions about an image. This ability is challenging but indispensable to achieve general VQA. One limita…
Knowledge GraphsQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)