paper-with-me

Papers

Interpretable Visual Question Answering Referring to Outside Knowledge

2023-03-08 · He Zhu, Ren Togo, Takahiro Ogawa, Miki Haseyama

We present a novel multimodal interpretable VQA model that can answer the question more accurately and generate diverse explanations. Although researchers have proposed several methods that can generate human-readable and fine-grained natural language sentences to explain a model's decision, these methods have focused solely on the information in the image. Ideally, the model should refer to various information inside and outside the image to correctly generate explanations, just as we use background knowledge daily. The proposed method incorporates information from outside knowledge and multiple image captions to increase the diversity of information available to the model. The contribution of this paper is to construct an interpretable visual question answering model using multimodal inputs to improve the rationality of generated results. Experimental results show that our model can outperform state-of-the-art methods regarding answer accuracy and explanation rationality.

📄 PDF Abstract BibTeX arXiv:2303.04388

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityImage CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Can Open Domain Question Answering Systems Answer Visual Knowledge Questions?

2022-02-09 · Jiawen Zhang, Abhijit Mishra, Avinesh P. V. S, Siddharth Patwardhan 외

The task of Outside Knowledge Visual Question Answering (OKVQA) requires an automatic system to answer natural language questions about pictures and images using external knowledge. We observe that many visual questions,…

Open-Domain Question AnsweringQuestion AnsweringQuestion RewritingVisual Question Answering+1

Object-centric Video Question Answering with Visual Grounding and Referring

2025-07-25 · Haochen Wang, Qirui Chen, Cilin Yan, Jiayin Cai 외 arxiv

Video Large Language Models (VideoLLMs) have recently demonstrated remarkable progress in general video understanding. However, existing models primarily focus on high-level comprehension and are limited to text-only res…

Video Question AnsweringObject SegmentationVisual Grounding

Passage Retrieval for Outside-Knowledge Visual Question Answering

2021-05-09 · Chen Qu, Hamed Zamani, Liu Yang, W. Bruce Croft 외

In this work, we address multi-modal information needs that contain text questions and images by focusing on passage retrieval for outside-knowledge visual question answering. This task requires access to outside knowled…

Image CaptioningObjectPassage RetrievalQuestion Answering+3

Pre-Training Multi-Modal Dense Retrievers for Outside-Knowledge Visual Question Answering

2023-06-28 · Alireza Salemi, Mahta Rafiee, Hamed Zamani

This paper studies a category of visual question answering tasks, in which accessing external knowledge is necessary for answering the questions. This category is called outside-knowledge visual question answering (OK-VQ…

Passage RetrievalQuestion AnsweringRetrievalVisual Question Answering+1

Entity-Focused Dense Passage Retrieval for Outside-Knowledge Visual Question Answering

2022-10-18 · Jialin Wu, Raymond J. Mooney

Most Outside-Knowledge Visual Question Answering (OK-VQA) systems employ a two-stage framework that first retrieves external knowledge given the visual question and then predicts the answer based on the retrieved content…

Passage RetrievalQuestion AnsweringRetrievalVisual Question Answering+1