paper-with-me

Papers

ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named Entities

2022-07-11 · SIGIR 2022 7 · Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne, Romaric Besançon, Jose G Moreno, Jesús Lovón Melgarejo

Whether to retrieve, answer, translate, or reason, multimodality opens up new challenges and perspectives. In this context, we are interested in answering questions about named entities grounded in a visual context using a Knowledge Base (KB). To benchmark this task, called KVQAE (Knowledge-based Visual Question Answering about named Entities), we provide ViQuAE, a dataset of 3.7K questions paired with images. This is the first KVQAE dataset to cover a wide range of entity types (e.g. persons, landmarks, and products). The dataset is annotated using a semi-automatic method. We also propose a KB composed of 1.5M Wikipedia articles paired with images. To set a baseline on the benchmark, we address KVQAE as a two-stage problem: Information Retrieval and Reading Comprehension, with both zero-and few-shot learning methods. The experiments empirically demonstrate the difficulty of the task, especially when questions are not about persons. This work paves the way for better multimodal entity representations and question answering. The dataset, KB, code, and semi-automatic annotation pipeline are freely available at https://github.com/PaulLerner/ViQuAE.

📄 PDF Abstract BibTeX

Code (1)

paullerner/viquae pytorch

Tasks

ArticlesFew-Shot LearningInformation RetrievalQuestion AnsweringReading ComprehensionRetrievalVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named Entities

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Whether to retrieve, answer, translate or reason, multimodality opens up new challenges and perspectives. In this context, we are interested in Knowledge-based Visual Question Answering about named Entities (KVQAE). To b…

ArticlesFace RecognitionImage RetrievalInformation Retrieval+6

Un jeu de données pour répondre à des questions visuelles à propos d’entités nommées en utilisant des bases de connaissances (ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named Entities)

2022-06-01 · JEP/TALN/RECITAL 2022 6 · Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne 외

Dans le contexte général des traitements multimodaux, nous nous intéressons à la tâche de réponse à des questions visuelles à propos d’entités nommées en utilisant des bases de connaissances (KVQAE). Nous mettons à dispo…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Cross-modal Retrieval for Knowledge-based Visual Question Answering

2024-01-11 · Paul Lerner, Olivier Ferret, Camille Guinaudeau

Knowledge-based Visual Question Answering about Named Entities is a challenging task that requires retrieving information from a multimodal Knowledge Base. Named entities have diverse visual representations and are there…

Cross-Modal RetrievalQuestion AnsweringRetrievalVisual Question Answering+1

Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering

2026-07-28 · Noor Islam S. Mohammad, Uluğ Bayazıt arxiv

Knowledge-intensive multimodal question answering (KI-MMQA) sits at the intersection of three expensive primitives: long visual token sequences, dense retrieval over large external corpora, and full cross-modal fusion. E…

Question Answering

A Dataset and Baselines for Visual Question Answering on Art

2020-08-28 · Noa Garcia, Chentao Ye, Zihua Liu, Qingtao Hu 외

Answering questions related to art pieces (paintings) is a difficult task, as it implies the understanding of not only the visual information that is shown in the picture, but also the contextual knowledge that is acquir…

Question AnsweringQuestion GenerationQuestion-GenerationVisual Question Answering+1