paper-with-me

홈 › Papers

Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question Answering

2018-11-01 · NeurIPS 2018 12 · Medhini Narasimhan, Svetlana Lazebnik, Alexander G. Schwing

Accurately answering a question about a given image requires combining observations with general knowledge. While this is effortless for humans, reasoning with general knowledge remains an algorithmic challenge. To advance research in this direction a novel fact-based' visual question answering (FVQA) task has been introduced recently along with a large set of curated facts which link two entities, i.e., two possible answers, via a relation. Given a question-image pair, deep network techniques have been employed to successively reduce the large set of facts until one of the two entities of the final remaining fact is predicted as the answer. We observe that a successive process which considers one fact at a time to form a local decision is sub-optimal. Instead, we develop an entity graph and use a graph convolutional network to reason' about the correct answer by jointly considering all entities. We show on the challenging FVQA dataset that this leads to an improvement in accuracy of around 7% compared to the state of the art.

📄 PDF Abstract BibTeX arXiv:1811.00538

Code (0)

등록된 구현이 없습니다.

Tasks

Factual Visual Question AnsweringGeneral KnowledgeQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Iterative Visual Reasoning Beyond Convolutions

2018-03-29 · CVPR 2018 6 · Xinlei Chen, Li-Jia Li, Li Fei-Fei, Abhinav Gupta

We present a novel framework for iterative visual reasoning. Our framework goes beyond current recognition systems that lack the capability to reason beyond stack of convolutions. The framework consists of two core modul…

Visual Reasoning

Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering

2020-06-16 · Zihao Zhu, Jing Yu, Yujing Wang, Yajing Sun 외

Fact-based Visual Question Answering (FVQA) requires external knowledge beyond visible content to answer questions about an image, which is challenging but indispensable to achieve general VQA. One limitation of existing…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models

2026-05-26 · Yifan Jiang, Dae Yon Hwang, Jesse C. Cresswell, Freda Shi arxiv

Chart question-answering (QA) benchmarks aim to pose questions that require visual reasoning to correctly answer, but models can often reach solutions through shortcuts or prior familiarity with a chart based on their ow…

Visual Reasoning

Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models

2025-11-25 · Shamima Hossain arxiv

Visual Language Models (VLMs) are powerful generative tools but often produce factually inaccurate outputs due to a lack of robust reasoning capabilities. While extensive research has been conducted on integrating extern…

Knowledge Graphs

Cross-modal Knowledge Reasoning for Knowledge-based Visual Question Answering

2020-08-31 · Jing Yu, Zihao Zhu, Yujing Wang, Weifeng Zhang 외

Knowledge-based Visual Question Answering (KVQA) requires external knowledge beyond the visible content to answer questions about an image. This ability is challenging but indispensable to achieve general VQA. One limita…

Knowledge GraphsQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)