VQA Therapy: Exploring Answer Differences by Visually Grounding Answers
Visual question answering is a task of predicting the answer to a question about an image. Given that different people can provide different answers to a visual question, we aim to better understand why with answer groundings. We introduce the first dataset that visually grounds each unique answer to each visual question, which we call VQAAnswerTherapy. We then propose two novel problems of predicting whether a visual question has a single answer grounding and localizing all answer groundings. We benchmark modern algorithms for these novel problems to show where they succeed and struggle. The dataset and evaluation server can be found publicly at https://vizwiz.org/tasks-and-datasets/vqa-answer-therapy/.
Code (1)
Tasks
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Grounding Answers for Visual Questions Asked by Visually Impaired People
Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQAGrounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual imp…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Grounding Answers for Visual Questions Asked by Visually Impaired People
Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQA-Grounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual im…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Accounting for Focus Ambiguity in Visual Questions
No existing work on visual question answering explicitly accounts for ambiguity regarding where the content described in the question is located in the image. To fill this gap, we introduce VQ-FocusAmbiguity, the first V…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA
High benchmark accuracy does not guarantee genuine use of visual evidence. We study this problem in traffic accident Video Question Answering (VideoQA), where correct answers should depend on scene-specific visual eviden…
Video Question AnsweringVisual GroundingCan I Trust Your Answer? Visually Grounded Video Question Answering
We study visually grounded VideoQA in response to the emerging trends of utilizing pretraining techniques for video-language understanding. Specifically, by forcing vision-language models (VLMs) to answer questions and s…
Grounded Video Question AnsweringQuestion AnsweringVideo GroundingVideo Question Answering+1