paper-with-me

홈 › Papers

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning

2025-05-22 · Mingxin Huang, Yongxin Shi, Dezhi Peng, Songxuan Lai, Zecheng Xie, Lianwen Jin

Recent advancements in multimodal slow-thinking systems have demonstrated remarkable performance across diverse visual reasoning tasks. However, their capabilities in text-rich image reasoning tasks remain understudied due to the lack of a systematic benchmark. To address this gap, we propose OCR-Reasoning, a comprehensive benchmark designed to systematically assess Multimodal Large Language Models on text-rich image reasoning tasks. The benchmark comprises 1,069 human-annotated examples spanning 6 core reasoning abilities and 18 practical reasoning tasks in text-rich visual scenarios. Furthermore, unlike other text-rich image understanding benchmarks that only annotate the final answers, OCR-Reasoning also annotates the reasoning process simultaneously. With the annotated reasoning process and the final answers, OCR-Reasoning evaluates not only the final answers generated by models but also their reasoning processes, enabling a holistic analysis of their problem-solving abilities. Leveraging this benchmark, we conducted a comprehensive evaluation of state-of-the-art MLLMs. Our results demonstrate the limitations of existing methodologies. Notably, even state-of-the-art MLLMs exhibit substantial difficulties, with none achieving accuracy surpassing 50\% across OCR-Reasoning, indicating that the challenges of text-rich image reasoning are an urgent issue to be addressed. The benchmark and evaluation scripts are available at https://github.com/SCUT-DLVCLab/OCR-Reasoning.

📄 PDF Abstract BibTeX arXiv:2505.17163

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character Recognition (OCR)Visual Reasoning

Similar Papers 제목 키워드 기반

MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models

2025-06-05 · Gio Paik, Geewook Kim, Jinbae Im

This paper introduces MMRefine, a MultiModal Refinement benchmark designed to evaluate the error refinement capabilities of Multimodal Large Language Models (MLLMs). As the emphasis shifts toward enhancing reasoning duri…

Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models

2023-12-29 · Yuqing Wang, Yun Zhao

The burgeoning interest in Multimodal Large Language Models (MLLMs), such as OpenAI's GPT-4V(ision), has significantly impacted both academic and industrial realms. These models enhance Large Language Models (LLMs) with …

HellaSwag

VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity

2025-03-14 · Jing Bi, Junjia Guo, Susan Liang, Guangyu Sun 외

Visual reasoning is central to human cognition, enabling individuals to interpret and abstractly understand their environment. Although recent Multimodal Large Language Models (MLLMs) have demonstrated impressive perform…

BenchmarkingDecision MakingMultimodal ReasoningVisual Reasoning

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

2025-10-27 · Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He 외 arxiv

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core chall…

Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning

2026-02-01 · Meng Luo, Bobo Li, Shanqing Xu, Shize Zhang 외 arxiv

Despite rapid progress in multimodal large language models (MLLMs), their capability for deep emotional understanding remains limited. We argue that genuine affective intelligence requires explicit modeling of Theory of …

Reinforcement Learning