A Dataset for Multimodal Question Answering in the Cultural Heritage Domain
Multimodal question answering in the cultural heritage domain allows visitors to ask questions in a more natural way and thus provides better user experiences with cultural objects while visiting a museum, landmark or any other historical site. In this paper, we introduce the construction of a golden standard dataset that will aid research of multimodal question answering in the cultural heritage domain. The dataset, which will be soon released to the public, contains multimodal content including images of typical artworks from the fascinating old-Egyptian Amarna period, related image-containing documents of the artworks and over 800 multimodal queries integrating visual and textual questions. The multimodal questions and related documents are all in English. The multimodal questions are linked to relevant paragraphs in the related documents that contain the answer to the multimodal query.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringSpeech RecognitionVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China
We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild…
Visual Question AnsweringVaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific data. We introduce VaseVQA, a benchmark of…
Visual Question AnsweringReinforcement LearningFoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture
Food is a rich and varied dimension of cultural heritage, crucial to both individuals and social groups. To bridge the gap in the literature on the often-overlooked regional diversity in this domain, we introduce FoodieQ…
DiversityMultiple-choiceQuestion AnsweringVisual Question Answering (VQA)Is GPT-3 all you need for Visual Question Answering in Cultural Heritage?
The use of Deep Learning and Computer Vision in the Cultural Heritage domain is becoming highly relevant in the last few years with lots of applications about audio smart guides, interactive museums and augmented reality…
AllQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Increasing the Difficulty of Automatically Generated Questions via Reinforcement Learning with Synthetic Preference
As the cultural heritage sector increasingly adopts technologies like Retrieval-Augmented Generation (RAG) to provide more personalised search experiences and enable conversations with collections data, the demand for sp…
Machine Reading ComprehensionQuestion AnsweringRAGReading Comprehension+2