paper-with-me

Papers

Visual7W: Grounded Question Answering in Images

2015-11-11 · CVPR 2016 6 · Yuke Zhu, Oliver Groth, Michael Bernstein, Li Fei-Fei

We have seen great progress in basic perceptual tasks such as object recognition and detection. However, AI models still fail to match humans in high-level vision tasks due to the lack of capacities for deeper reasoning. Recently the new task of visual question answering (QA) has been proposed to evaluate a model's capacity for deep image understanding. Previous works have established a loose, global association between QA sentences and images. However, many questions and answers, in practice, relate to local regions in the images. We establish a semantic link between textual descriptions and image regions by object-level grounding. It enables a new type of QA with visual answers, in addition to textual answers used in previous work. We study the visual QA tasks in a grounded setting with a large collection of 7W multiple-choice QA pairs. Furthermore, we evaluate human performance and several baseline models on the QA tasks. Finally, we propose a novel LSTM model with spatial attention to tackle the 7W QA tasks.

📄 PDF Abstract BibTeX arXiv:1511.03416

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple-choiceMultiple Choice Question Answering (MCQA)Object RecognitionQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection

2025-12-15 · Juil Koo, Daehyeon Choi, Sangwoo Youn, Phillip Y. Lee 외 arxiv

Vision Language Models (VLMs) excel at visual question answering (VQA) but remain limited to snapshot vision, reasoning from static images. In contrast, embodied agents require ambulatory vision, actively moving to obtai…

Visual Question Answering

Can Open Domain Question Answering Systems Answer Visual Knowledge Questions?

2022-02-09 · Jiawen Zhang, Abhijit Mishra, Avinesh P. V. S, Siddharth Patwardhan 외

The task of Outside Knowledge Visual Question Answering (OKVQA) requires an automatic system to answer natural language questions about pictures and images using external knowledge. We observe that many visual questions,…

Open-Domain Question AnsweringQuestion AnsweringQuestion RewritingVisual Question Answering+1

Multimodal Inverse Cloze Task for Knowledge-based Visual Question Answering

2023-01-11 · Paul Lerner, Olivier Ferret, Camille Guinaudeau

We present a new pre-training method, Multimodal Inverse Cloze Task, for Knowledge-based Visual Question Answering about named Entities (KVQAE). KVQAE is a recently introduced task that consists in answering questions ab…

Question AnsweringReading ComprehensionRetrievalSentence+2

WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark from Wikipedia and Wikidata

2026-05-20 · Basel Shbita, Pengyuan Li, Anna Lisa Gentile arxiv

Visual Question Answering (VQA) benchmarks have largely emphasized perception-based tasks that can be solved from visual content alone. In contrast, many real-world scenarios require external knowledge that is not direct…

Visual Question Answering

ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named Entities

2022-07-11 · SIGIR 2022 7 · Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne 외

Whether to retrieve, answer, translate, or reason, multimodality opens up new challenges and perspectives. In this context, we are interested in answering questions about named entities grounded in a visual context using…

ArticlesFew-Shot LearningInformation RetrievalQuestion Answering+4