paper-with-me

홈 › Papers

Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?

2025-05-19 · Haibin He, Maoyuan Ye, Jing Zhang, Xiantao Cai, Juhua Liu, Bo Du, DaCheng Tao

Large Multimodal Models (LMMs) have become increasingly versatile, accompanied by impressive Optical Character Recognition (OCR) related capabilities. Existing OCR-related benchmarks emphasize evaluating LMMs' abilities of relatively simple visual question answering, visual-text parsing, etc. However, the extent to which LMMs can deal with complex logical reasoning problems based on OCR cues is relatively unexplored. To this end, we introduce the Reasoning-OCR benchmark, which challenges LMMs to solve complex reasoning problems based on the cues that can be extracted from rich visual-text. Reasoning-OCR covers six visual scenarios and encompasses 150 meticulously designed questions categorized into six reasoning challenges. Additionally, Reasoning-OCR minimizes the impact of field-specialized knowledge. Our evaluation offers some insights for proprietary and open-source LMMs in different reasoning challenges, underscoring the urgent to improve the reasoning performance. We hope Reasoning-OCR can inspire and facilitate future research on enhancing complex reasoning ability based on OCR cues. Reasoning-OCR is publicly available at https://github.com/Hxyz-123/ReasoningOCR.

📄 PDF Abstract BibTeX arXiv:2505.12766

Code (1)

hxyz-123/reasoningocr 공식 구현

Tasks

Logical ReasoningOptical Character RecognitionOptical Character Recognition (OCR)Question AnsweringVisual Question Answering

Similar Papers 제목 키워드 기반

Can Multimodal Large Language Model Think Analogically?

2024-11-02 · Diandian Guo, Cong Cao, Fangfang Yuan, Dakui Wang 외

Analogical reasoning, particularly in multimodal contexts, is the foundation of human perception and creativity. Multimodal Large Language Model (MLLM) has recently sparked considerable discussion due to its emergent cap…

Language ModelingLanguage ModellingLarge Language Modelmodel+1

MuSLR: Multimodal Symbolic Logical Reasoning

2025-09-30 · Jundong Xu, Hao Fei, Yuhui Zhang, Liangming Pan 외 arxiv

Multimodal symbolic logical reasoning, which aims to deduce new facts from multimodal input via formal logic, is critical in high-stakes applications such as autonomous driving and medical diagnosis, as its rigorous, det…

Autonomous DrivingMedical DiagnosisLogical ReasoningFormal Logic

R1-Onevision:An Open-Source Multimodal Large Language Model Capable of Deep Reasoning

2025-02-24 · ongoing 2025 2 · Yi Yang*, Xiaoxuan He*, Hongkun Pan*, Xiyan Jiang 외

R1-OneVision is a versatile multimodal reasoning large model, designed to tackle complex visual reasoning tasks. It seamlessly integrates visual and textual data to offer precise interpretations of multimodal information…

Language ModelingLanguage ModellingLarge Language ModelLogical Reasoning+3

LoRA: A Logical Reasoning Augmented Dataset for Visual Question Answering

2023-09-26 · NeurIPS 2023 11

The capacity to reason logically is a hallmark of human cognition. Humans excel at integrating multimodal information for locigal reasoning, as exemplified by the Visual Question Answering (VQA) task, which is a challeng…

Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models

2024-03-26 · Zhenyu Pan, Haozheng Luo, Manling Li, Han Liu

We present a Chain-of-Action (CoA) framework for multimodal and retrieval-augmented Question-Answering (QA). Compared to the literature, CoA overcomes two major challenges of current QA applications: (i) unfaithful hallu…

HallucinationInformation RetrievalQuestion AnsweringRetrieval