Inverse Visual Question Answering with Multi-Level Attentions
In this paper, we propose a novel deep multi-level attention model to address inverse visual question answering. The proposed model generates regional visual and semantic features at the object level and then enhances them with the answer cue by using attention mechanisms. Two levels of multiple attentions are employed in the model, including the dual attention at the partial question encoding step and the dynamic attention at the next question word generation step. We evaluate the proposed model on the VQA V1 dataset. It demonstrates state-of-the-art performance in terms of multiple commonly used metrics.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Multimodal Inverse Cloze Task for Knowledge-based Visual Question Answering
We present a new pre-training method, Multimodal Inverse Cloze Task, for Knowledge-based Visual Question Answering about named Entities (KVQAE). KVQAE is a recently introduced task that consists in answering questions ab…
Question AnsweringReading ComprehensionRetrievalSentence+2iVQA: Inverse Visual Question Answering
We propose the inverse problem of Visual question answering (iVQA), and explore its suitability as a benchmark for visuo-linguistic understanding. The iVQA task is to generate a question that corresponds to a given image…
Question AnsweringQuestion GenerationQuestion-GenerationVisual Question Answering+1Multi-Level Attention Networks for Visual Question Answering
Inspired by the recent success of text-based question answering, visual question answering (VQA) is proposed to automatically answer natural language questions with the reference to a given image. Compared with text-base…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Multilevel Hierarchical Network with Multiscale Sampling for Video Question Answering
Video question answering (VideoQA) is challenging given its multimodal combination of visual understanding and natural language processing. While most existing approaches ignore the visual appearance-motion information a…
multimodal interactionQuestion AnsweringVideo Question AnsweringVisual ReasoningVideo Question Answering via Attribute-Augmented Attention Network Learning
Video Question Answering is a challenging problem in visual information retrieval, which provides the answer to the referenced video content according to the question. However, the existing visual question answering appr…
AttributeInformation RetrievalMultiple-choiceQuestion Answering+5