paper-with-me

홈 › Papers

A Dataset for Multimodal Question Answering in the Cultural Heritage Domain

2016-12-01 · WS 2016 12 · Shurong Sheng, Luc van Gool, Marie-Francine Moens

Multimodal question answering in the cultural heritage domain allows visitors to ask questions in a more natural way and thus provides better user experiences with cultural objects while visiting a museum, landmark or any other historical site. In this paper, we introduce the construction of a golden standard dataset that will aid research of multimodal question answering in the cultural heritage domain. The dataset, which will be soon released to the public, contains multimodal content including images of typical artworks from the fascinating old-Egyptian Amarna period, related image-containing documents of the artworks and over 800 multimodal queries integrating visual and textual questions. The multimodal questions and related documents are all in English. The multimodal questions are linked to relevant paragraphs in the related documents that contain the answer to the multimodal query.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringSpeech RecognitionVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

2026-06-08 · Yi Zhang, Bolei Ma, Yong Cao, Chengyan Wu 외 arxiv

We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild…

Visual Question Answering

VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery

2025-09-21 · Jinchao Ge, Tengfei Cheng, Biao Wu, Zeyu Zhang 외 arxiv

Understanding cultural heritage artifacts such as ancient Greek pottery requires expert-level reasoning that remains challenging for current MLLMs due to limited domain-specific data. We introduce VaseVQA, a benchmark of…

Visual Question AnsweringReinforcement Learning

FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food Culture

2024-06-16 · Wenyan Li, Xinyu Zhang, Jiaang Li, Qiwei Peng 외

Food is a rich and varied dimension of cultural heritage, crucial to both individuals and social groups. To bridge the gap in the literature on the often-overlooked regional diversity in this domain, we introduce FoodieQ…

DiversityMultiple-choiceQuestion AnsweringVisual Question Answering (VQA)

Is GPT-3 all you need for Visual Question Answering in Cultural Heritage?

2022-07-25 · Pietro Bongini, Federico Becattini, Alberto del Bimbo

The use of Deep Learning and Computer Vision in the Cultural Heritage domain is becoming highly relevant in the last few years with lots of applications about audio smart guides, interactive museums and augmented reality…

AllQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Increasing the Difficulty of Automatically Generated Questions via Reinforcement Learning with Synthetic Preference

2024-10-10 · William Thorne, Ambrose Robinson, Bohua Peng, Chenghua Lin 외

As the cultural heritage sector increasingly adopts technologies like Retrieval-Augmented Generation (RAG) to provide more personalised search experiences and enable conversations with collections data, the demand for sp…

Machine Reading ComprehensionQuestion AnsweringRAGReading Comprehension+2