paper-with-me

홈 › Papers

MultiOCR-QA: Dataset for Evaluating Robustness of LLMs in Question Answering on Multilingual OCR Texts

2025-02-24 · Bhawna Piryani, Jamshid Mozafari, Abdelrahman Abdallah, Antoine Doucet, Adam Jatowt

Optical Character Recognition (OCR) plays a crucial role in digitizing historical and multilingual documents, yet OCR errors -- imperfect extraction of the text, including character insertion, deletion and permutation -- can significantly impact downstream tasks like question-answering (QA). In this work, we introduce a multilingual QA dataset MultiOCR-QA, designed to analyze the effects of OCR noise on QA systems' performance. The MultiOCR-QA dataset comprises 60K question-answer pairs covering three languages, English, French, and German. The dataset is curated from OCR-ed old documents, allowing for the evaluation of OCR-induced challenges on question answering. We evaluate MultiOCR-QA on various levels and types of OCR errors to access the robustness of LLMs in handling real-world digitization errors. Our findings show that QA systems are highly prone to OCR induced errors and exhibit performance degradation on noisy OCR text.

📄 PDF Abstract BibTeX arXiv:2502.16781

Code (1)

datascienceuibk/multiocr-qa 공식 구현

Tasks

Optical Character RecognitionOptical Character Recognition (OCR)Question Answering

Similar Papers 제목 키워드 기반

HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning

2025-02-17 · Xiaoyuan Li, Moxin Li, Rui Men, Yichang Zhang 외

Large language models (LLMs) have shown remarkable capabilities in commonsense reasoning; however, some variations in questions can trigger incorrect responses. Do these models truly understand commonsense knowledge, or …

HellaSwag

Evaluating the Performance and Robustness of LLMs in Materials Science Q&A and Property Predictions

2024-09-22 · Hongchen Wang, Kangming Li, Scott Ramsay, Yao Fehlis 외

Large Language Models (LLMs) have the potential to revolutionize scientific research, yet their robustness and reliability in domain-specific applications remain insufficiently explored. In this study, we evaluate the pe…

Band GapIn-Context LearningMultiple-choiceProperty Prediction+1

ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty

2024-12-28 · Qing Zong, Zhaowei Wang, Tianshi Zheng, Xiyu Ren 외

The rapid development of LLMs has sparked extensive research into their factual knowledge. Current works claim that LLMs fall short on questions requiring less frequent knowledge. However, their proof is incomplete since…

GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

2024-02-29 · Qintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong 외

Large language models (LLMs) have achieved impressive performance across various mathematical reasoning benchmarks. However, there are increasing debates regarding whether these models truly understand and apply mathemat…

GSM8KMathMathematical ReasoningMath Word Problem Solving

BioPulse-QA: A Dynamic Biomedical Question-Answering Benchmark for Evaluating Factuality, Robustness, and Bias in Large Language Models

2026-01-19 · Kriti Bhattarai, Vipina K. Keloth, Donald Wright, Andrew Loza 외 arxiv

Objective: Large language models (LLMs) are increasingly applied in biomedical settings, and existing benchmark datasets have played an important role in supporting model development and evaluation. However, these benchm…

Question Answering