paper-with-me

홈 › Papers

Towards VQA Models That Can Read

2019-04-18 · CVPR 2019 6 · Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, Marcus Rohrbach

Studies have shown that a dominant class of questions asked by visually impaired users on images of their surroundings involves reading text in the image. But today's VQA models can not read! Our paper takes a first step towards addressing this problem. First, we introduce a new "TextVQA" dataset to facilitate progress on this important problem. Existing datasets either have a small proportion of questions about text (e.g., the VQA dataset) or are too small (e.g., the VizWiz dataset). TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Second, we introduce a novel model architecture that reads text in the image, reasons about it in the context of the image and the question, and predicts an answer which might be a deduction based on the text and the image or composed of the strings found in the image. Consequently, we call our approach Look, Read, Reason & Answer (LoRRA). We show that LoRRA outperforms existing state-of-the-art VQA models on our TextVQA dataset. We find that the gap between human performance and machine performance is significantly larger on TextVQA than on VQA 2.0, suggesting that TextVQA is well-suited to benchmark progress along directions complementary to VQA 2.0.

📄 PDF Abstract BibTeX arXiv:1904.08920

Code (7)

facebookresearch/pythia 공식 구현 pytorch
ZephyrZhuQi/ssbaseline pytorch
allenai/pythia pytorch
facebookresearch/mmf pytorch
jackroos/pythia pytorch
ronghanghu/pythia pytorch
zwxalgorithm/pythia pytorch

Tasks

TextVQAVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

On Understanding the Relation between Expert Annotations of Text Readability and Target Reader Comprehension

2019-08-01 · WS 2019 8 · Sowmya Vajjala, Ivana Lucic

Automatic readability assessment aims to ensure that readers read texts that they can comprehend. However, computational models are typically trained on texts created from the perspective of the text writer, not the targ…

Question GenerationQuestion-GenerationReading ComprehensionRelation

Generating Summaries with Controllable Readability Levels

2023-10-16 · Leonardo F. R. Ribeiro, Mohit Bansal, Markus Dreyer

Readability refers to how easily a reader can understand a written text. Several factors affect the readability level, such as the complexity of the text, its subject matter, and the reader's background knowledge. Genera…

Lay SummarizationNews SummarizationText Generation

My Turn To Read: An Interleaved E-book Reading Tool for Developing and Struggling Readers

2019-07-01 · ACL 2019 7 · Nitin Madnani, Beata Beigman Klebanov, Anastassia Loukina, Binod Gyawali 외

Literacy is crucial for functioning in modern society. It underpins everything from educational attainment and employment opportunities to health outcomes. We describe My Turn To Read, an app that uses interleaved readin…

Eye Tracking Based Cognitive Evaluation of Automatic Readability Assessment Measures

2025-02-16 · Keren Gruteke Klein, Shachar Frenkel, Omer Shubi, Yevgeni Berzak

Automated text readability prediction is widely used in many real-world scenarios. Over the past century, such measures have primarily been developed and evaluated on reading comprehension outcomes and on human annotatio…

AllReading ComprehensionText Simplification

The LetsRead Corpus of Portuguese Children Reading Aloud for Performance Evaluation

2016-05-01 · LREC 2016 5 · Jorge Proen{\c{c}}a, Dirce Celorico, C, Sara eias 외

This paper introduces the LetsRead Corpus of European Portuguese read speech from 6 to 10 years old children. The motivation for the creation of this corpus stems from the inexistence of databases with recordings of read…