paper-with-me

홈 › Papers

ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models

2025-05-25 · Benjamin Clavié, Florian Brand

Recent advancements in Large Vision-Language Models (VLMs), have greatly enhanced their capability to jointly process text and images. However, despite extensive benchmarks evaluating visual comprehension (e.g., diagrams, color schemes, OCR tasks...), there is limited assessment of VLMs' ability to read and reason about text-rich images effectively. To fill this gap, we introduce ReadBench, a multimodal benchmark specifically designed to evaluate the reading comprehension capabilities of VLMs. ReadBench transposes contexts from established text-only benchmarks into images of text while keeping textual prompts and questions intact. Evaluating leading VLMs with ReadBench, we find minimal-but-present performance degradation on short, text-image inputs, while performance sharply declines for longer, multi-page contexts. Our experiments further reveal that text resolution has negligible effects on multimodal performance. These findings highlight needed improvements in VLMs, particularly their reasoning over visually presented extensive textual content, a capability critical for practical applications. ReadBench is available at https://github.com/answerdotai/ReadBench .

📄 PDF Abstract BibTeX arXiv:2505.19091

Code (1)

answerdotai/readbench 공식 구현

Tasks

Optical Character Recognition (OCR)Reading Comprehension

Similar Papers 제목 키워드 기반

Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench

2025-10-30 · Fenfen Lin, Yesheng Liu, Haiyu Xu, Chen Yue 외 arxiv

Reading measurement instruments is effortless for humans and requires relatively little domain expertise, yet it remains surprisingly challenging for current vision-language models (VLMs) as we find in preliminary evalua…

ICDAR 2023 Video Text Reading Competition for Dense and Small Text

2023-04-10 · Weijia Wu, Yuzhong Zhao, Zhuang Li, Jiahong Li 외

Recently, video text detection, tracking, and recognition in natural scenes are becoming very popular in the computer vision community. However, most existing algorithms and benchmarks focus on common text cases (e.g., n…

Task 2Text DetectionText Spottingvalid

Digital Comprehensibility Assessment of Simplified Texts among Persons with Intellectual Disabilities

2024-02-20 · Andreas Säuberli, Franz Holzknecht, Patrick Haller, Silvana Deilen 외

Text simplification refers to the process of increasing the comprehensibility of texts. Automatic text simplification models are most commonly evaluated by experts or crowdworkers instead of the primary target groups of …

Multiple-choiceText Simplification

Measuring text readability with machine comprehension: a pilot study

2019-08-01 · WS 2019 8 · Marc Benzahra, Fran{\c{c}}ois Yvon

This article studies the relationship between text readability indice and automatic machine understanding systems. Our hypothesis is that the simpler a text is, the better it should be understood by a machine. We thus ex…

Reading Comprehension

Lip-reading with Densely Connected Temporal Convolutional Networks

2020-09-29 · Pingchuan Ma, Yujiang Wang, Jie Shen, Stavros Petridis 외

In this work, we present the Densely Connected Temporal Convolutional Network (DC-TCN) for lip-reading of isolated words. Although Temporal Convolutional Networks (TCN) have recently demonstrated great potential in many …

Lip Reading