Visual Question Answering v2.0
VQA v2.0
홈페이지 · 논문 366편
Visual Question Answering (VQA) v2.0 is a dataset containing open-ended questions about images. These questions require an understanding of vision, language and commonsense knowledge to answer. It is the second version of the VQA dataset. - 265,016 images (COCO and abstract scenes) - At least 3 questions (5.4 questions on average) per image - 10 ground truth answers per question - 3 plausible (but likely incorrect) answers per question - Automatic evaluation metric The [first version of the dataset](/dataset/visual-question-answering) was released in October 2015.
ImagesTexts