paper-with-me

홈 › Papers

DisasterVQA: A Visual Question Answering Benchmark Dataset for Disaster Scenes

2026-01-20 · Aisha Al-Mohannadi, Ayisha Firoz, Yin Yang, Muhammad Imran, Ferda Ofli arxiv

Social media imagery provides a low-latency source of situational information during natural and human-induced disasters, enabling rapid damage assessment and response. While Visual Question Answering (VQA) has shown strong performance in general-purpose domains, its suitability for the complex and safety-critical reasoning required in disaster response remains unclear. We introduce DisasterVQA, a benchmark dataset designed for perception and reasoning in crisis contexts. DisasterVQA consists of 1,395 real-world images and 4,405 expert-curated question-answer pairs spanning diverse events such as floods, wildfires, and earthquakes. Grounded in humanitarian frameworks including FEMA ESF and OCHA MIRA, the dataset includes binary, multiple-choice, and open-ended questions covering situational awareness and operational decision-making tasks. We benchmark seven state-of-the-art vision-language models and find performance variability across question types, disaster categories, regions, and humanitarian tasks. Although models achieve high accuracy on binary questions, they struggle with fine-grained quantitative reasoning, object counting, and context-sensitive interpretation, particularly for underrepresented disaster scenarios. DisasterVQA provides a challenging and practical benchmark to guide the development of more robust and operationally meaningful vision-language models for disaster response. The dataset is publicly available at https://doi.org/10.5281/zenodo.18267769.

📄 PDF Abstract BibTeX arXiv:2601.13839

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringObject Counting

Similar Papers 제목 키워드 기반

OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

2019-05-31 · CVPR 2019 6 · Kenneth Marino, Mohammad Rastegari, Ali Farhadi, Roozbeh Mottaghi

Visual Question Answering (VQA) in its ideal form lets us study reasoning in the joint space of vision and language and serves as a proxy for the AI task of scene understanding. However, most VQA benchmarks to date are f…

object-detectionObject DetectionQuestion AnsweringScene Understanding+2

Answer-Type Prediction for Visual Question Answering

2016-06-01 · CVPR 2016 6 · Kushal Kafle, Christopher Kanan

Recently, algorithms for object recognition and related tasks have become sufficiently proficient that new vision tasks can now be pursued. In this paper, we build a system capable of answering open-ended text-based ques…

Object RecognitionPredictionQuestion AnsweringType prediction+3

Multi-Image Visual Question Answering

2021-12-27 · Harsh Raj, Janhavi Dadhania, Akhilesh Bhardwaj, Prabuchandran KJ

While a lot of work has been done on developing models to tackle the problem of Visual Question Answering, the ability of these models to relate the question to the image features still remain less explored. We present a…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

EVJVQA Challenge: Multilingual Visual Question Answering

2023-02-23 · Ngan Luu-Thuy Nguyen, Nghia Hieu Nguyen, Duong T. D Vo, Khanh Quoc Tran 외

Visual Question Answering (VQA) is a challenging task of natural language processing (NLP) and computer vision (CV), attracting significant attention from researchers. English is a resource-rich language that has witness…

Language ModelingLanguage ModellingQuestion AnsweringVietnamese Multimodal Learning+3

ViQuAE, a Dataset for Knowledge-based Visual Question Answering about Named Entities

2022-07-11 · SIGIR 2022 7 · Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne 외

Whether to retrieve, answer, translate, or reason, multimodality opens up new challenges and perspectives. In this context, we are interested in answering questions about named entities grounded in a visual context using…

ArticlesFew-Shot LearningInformation RetrievalQuestion Answering+4