paper-with-me

TextVQA

홈페이지 · 논문 476편

TextVQA is a dataset to benchmark visual reasoning based on text in images. TextVQA requires models to read and reason about text in images to answer questions about them. Specifically, models need to incorporate a new modality of text present in the images and reason over it to answer TextVQA questions. Statistics * 28,408 images from OpenImages * 45,336 questions * 453,360 ground truth answers

ImagesTexts

벤치마크

Visual Question Answering (VQA) on TextVQA test-standard 결과 12개
Visual Question Answering on TextVQA test-standard 결과 2개
Visual Question Answering (VQA) on TextVQA 결과 1개