paper-with-me

홈 › Papers

VizWiz Grand Challenge: Answering Visual Questions from Blind People

2018-02-22 · CVPR 2018 6 · Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, Jeffrey P. Bigham

The study of algorithms to automatically answer visual questions currently is motivated by visual question answering (VQA) datasets constructed in artificial VQA settings. We propose VizWiz, the first goal-oriented VQA dataset arising from a natural VQA setting. VizWiz consists of over 31,000 visual questions originating from blind people who each took a picture using a mobile phone and recorded a spoken question about it, together with 10 crowdsourced answers per visual question. VizWiz differs from the many existing VQA datasets because (1) images are captured by blind photographers and so are often poor quality, (2) questions are spoken and so are more conversational, and (3) often visual questions cannot be answered. Evaluation of modern algorithms for answering visual questions and deciding if a visual question is answerable reveals that VizWiz is a challenging dataset. We introduce this dataset to encourage a larger community to develop more generalized algorithms that can assist blind people.

📄 PDF Abstract BibTeX arXiv:1802.08218

Code (1)

ghazaleh-mahmoodi/lxmert_compression pytorch

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Tell Me the Evidence? Dual Visual-Linguistic Interaction for Answer Grounding

2022-06-21 · Junwen Pan, Guanlin Chen, Yi Liu, Jiexiang Wang 외

Answer grounding aims to reveal the visual evidence for visual question answering (VQA), which entails highlighting relevant positions in the image when answering questions about images. Previous attempts typically tackl…

DecoderQuestion AnsweringVisual GroundingVisual Question Answering+1

Grounding Answers for Visual Questions Asked by Visually Impaired People

2022-06-20 · CVPR 2022 6 · Chongyan Chen; Samreen Anjum; Danna Gurari

Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQAGrounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual imp…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Grounding Answers for Visual Questions Asked by Visually Impaired People

2022-02-04 · CVPR 2022 1 · Chongyan Chen, Samreen Anjum, Danna Gurari

Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQA-Grounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual im…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Long-Form Answers to Visual Questions from Blind and Low Vision People

2024-08-12 · Mina Huh, Fangyuan Xu, Yi-Hao Peng, Chongyan Chen 외

Vision language models can now generate long-form answers to questions about images - long-form visual question answers (LFVQA). We contribute VizWiz-LF, a dataset of long-form answers to visual questions posed by blind …

FormVisual Question Answering (VQA)

Vision And Text Transformer For Predicting Answerability On Visual Question Answering

2026-09-15 · Tung Le, Huy Tien Nguyen, Le Minh Nguyen arxiv

Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question …

Visual Question Answering