VizWiz Grand Challenge: Answering Visual Questions from Blind People
The study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA dataset arising from a natural VQA setting. VizWiz consists of over 31,000 visual questions originating from blind people who each took a picture using a mobile phone and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. VizWiz differs from the many existing VQA datasets because (1) images are captured by blind photographers and so are often poor quality, (2) questions are spoken and so are more conversational, and (3) often visual questions cannot be answered. Evaluation of modern algorithms for answering visual questions and deciding if a visual question is answerable reveals that VizWiz is a challenging dataset. We introduce this dataset to encourage a larger community to develop more generalized algorithms that can assist blind people.
Code (1)
Tasks
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
Tell Me the Evidence? Dual Visual-Linguistic Interaction for Answer Grounding
Answer grounding aims to reveal the visual evidence for visual question answering (VQA), which entails highlighting relevant positions in the image when answering questions about images. Previous attempts typically tackl…
DecoderQuestion AnsweringVisual GroundingVisual Question Answering+1Grounding Answers for Visual Questions Asked by Visually Impaired People
Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQAGrounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual imp…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Grounding Answers for Visual Questions Asked by Visually Impaired People
Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQA-Grounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual im…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Long-Form Answers to Visual Questions from Blind and Low Vision People
Vision language models can now generate long-form answers to questions about images - long-form visual question answers (LFVQA). We contribute VizWiz-LF, a dataset of long-form answers to visual questions posed by blind …
FormVisual Question Answering (VQA)Vision And Text Transformer For Predicting Answerability On Visual Question Answering
Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question …
Visual Question Answering