paper-with-me

홈 › Papers

Context-VQA: Towards Context-Aware and Purposeful Visual Question Answering

2023-07-28 · Nandita Naik, Christopher Potts, Elisa Kreiss

Visual question answering (VQA) has the potential to make the Internet more accessible in an interactive way, allowing people who cannot see images to ask questions about them. However, multiple studies have shown that people who are blind or have low-vision prefer image explanations that incorporate the context in which an image appears, yet current VQA datasets focus on images in isolation. We argue that VQA models will not fully succeed at meeting people's needs unless they take context into account. To further motivate and analyze the distinction between different contexts, we introduce Context-VQA, a VQA dataset that pairs images with contexts, specifically types of websites (e.g., a shopping website). We find that the types of questions vary systematically across contexts. For example, images presented in a travel context garner 2 times more "Where?" questions, and images on social media and news garner 2.8 and 1.8 times more "Who?" questions than the average. We also find that context effects are especially important when participants can't see the image. These results demonstrate that context affects the types of questions asked and that VQA models should be context-sensitive to better meet people's needs, especially in accessibility settings.

📄 PDF Abstract BibTeX arXiv:2307.15745

Code (1)

nnaik39/context-vqa 공식 구현

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Travel 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos

2025-12-19 · Henghui Du, Chunjie Zhang, Xi Chen, Chang Zhou 외 arxiv

Long Video Question-Answering (LVQA) presents a significant challenge for Multi-modal Large Language Models (MLLMs) due to immense context and overloaded information, which could also lead to prohibitive memory consumpti…

CapWAP: Captioning with a Purpose

2020-11-09 · Adam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan H. Clark 외

The traditional image captioning task uses generic reference captions to provide textual information about images. Different user populations, however, will care about different visual aspects of images. In this paper, w…

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

CapWAP: Image Captioning with a Purpose

2020-11-01 · EMNLP 2020 11 · Adam Fisch, Kenton Lee, Ming-Wei Chang, Jonathan Clark 외

The traditional image captioning task uses generic reference captions to provide textual information about images. Different user populations, however, will care about different visual aspects of images. In this paper, w…

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Guiding Multimodal Large Language Models with Blind and Low Vision People Visual Questions for Proactive Visual Interpretations

2025-10-02 · Ricardo Gonzalez Penuela, Felipe Arias-Russi, Victor Capriles arxiv

Multimodal large language models (MLLMs) have been integrated into visual interpretation applications to support Blind and Low Vision (BLV) users because of their accuracy and ability to provide rich, human-like interpre…

GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering

2026-08-19 · Xinjin Li, Yudi Xia, Xi Zhao, Yiliu Xu 외 arxiv

Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence. Compared with conventional open-ended VQA, educational exampl…

Visual Question Answering