paper-with-me

GQA

홈페이지 · 논문 749편

The GQA dataset is a large-scale visual question answering dataset with real images from the Visual Genome dataset and balanced question-answer pairs. Each training and validation image is also associated with scene graph annotations describing the classes and attributes of those objects in the scene, and their pairwise relations. Along with the images and question-answer pairs, the GQA dataset provides two types of pre-extracted visual features for each image – convolutional grid features of size 7×7×2048 extracted from a ResNet-101 network trained on ImageNet, and object detection features of size Ndet×2048 (where Ndet is the number of detected objects in each image with a maximum of 100 per image) from a Faster R-CNN detector. Source: Language-Conditioned Graph Networks for Relational Reasoning Image Source: https://arxiv.org/pdf/1902.09506.pdf

ImagesTexts

벤치마크

Visual Question Answering (VQA) on GQA Test2019 결과 127개
Visual Question Answering (VQA) on GQA test-dev 결과 17개
Visual Question Answering (VQA) on GQA test-std 결과 7개
Object Detection on GQA 결과 5개
Scene Graph Generation on GQA 결과 3개
Graph Question Answering on GQA 결과 2개
Visual Question Answering on GQA 결과 2개
Visual Question Answering (VQA) on GQA 결과 2개