ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named Entities
Whether to retrieve, answer, translate or reason, multimodality opens up new challenges and perspectives. In this context, we are interested in Knowledge-based Visual Question Answering about named Entities (KVQAE). To benchmark the task, we provide ViQuAE, a dataset of 3.7K questions paired with images. This is the first KVQAE dataset to cover a wide range of entity types (e.g. persons, landmarks, products). The dataset is annotated using a semi-automatic method that could be extended to larger data scales. We also propose a Knowledge Base (KB) based on Wikipedia composed of 1.5M articles paired with images. To set a baseline on the benchmark, we address KVQAE as a two-stage problem: Information Retrieval (IR) and Reading Comprehension (RC). IR is carried out with a combination of face recognition, image retrieval and text retrieval while RC is purely text-based. The experiments empirically demonstrate the difficulty of the task. This work paves the way towards better multimodal entity representations and question answering. The dataset, KB and code will be available at https://github.com/Anonymous/ViQuAE.
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesFace RecognitionImage RetrievalInformation RetrievalQuestion AnsweringReading ComprehensionRetrievalText RetrievalVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named Entities
Whether to retrieve, answer, translate, or reason, multimodality opens up new challenges and perspectives. In this context, we are interested in answering questions about named entities grounded in a visual context using…
ArticlesFew-Shot LearningInformation RetrievalQuestion Answering+4Un jeu de données pour répondre à des questions visuelles à propos d’entités nommées en utilisant des bases de connaissances (ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named Entities)
Dans le contexte général des traitements multimodaux, nous nous intéressons à la tâche de réponse à des questions visuelles à propos d’entités nommées en utilisant des bases de connaissances (KVQAE). Nous mettons à dispo…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Cross-modal Retrieval for Knowledge-based Visual Question Answering
Knowledge-based Visual Question Answering about Named Entities is a challenging task that requires retrieving information from a multimodal Knowledge Base. Named entities have diverse visual representations and are there…
Cross-Modal RetrievalQuestion AnsweringRetrievalVisual Question Answering+1Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering
Knowledge-intensive multimodal question answering (KI-MMQA) sits at the intersection of three expensive primitives: long visual token sequences, dense retrieval over large external corpora, and full cross-modal fusion. E…
Question AnsweringA Dataset and Baselines for Visual Question Answering on Art
Answering questions related to art pieces (paintings) is a difficult task, as it implies the understanding of not only the visual information that is shown in the picture, but also the contextual knowledge that is acquir…
Question AnsweringQuestion GenerationQuestion-GenerationVisual Question Answering+1