paper-with-me

홈 › Papers

Evaluating LLMs for Historical Document OCR: A Methodological Framework for Digital Humanities

2025-10-08 · Maria Levchenko arxiv

Digital humanities scholars increasingly use Large Language Models for historical document digitization, yet lack appropriate evaluation frameworks for LLM-based OCR. Traditional metrics fail to capture temporal biases and period-specific errors crucial for historical corpus creation. We present an evaluation methodology for LLM-based historical OCR, addressing contamination risks and systematic biases in diplomatic transcription. Using 18th-century Russian Civil font texts, we introduce novel metrics including Historical Character Preservation Rate (HCPR) and Archaic Insertion Rate (AIR), alongside protocols for contamination control and stability testing. We evaluate 12 multimodal LLMs, finding that Gemini and Qwen models outperform traditional OCR while exhibiting over-historicization: inserting archaic characters from incorrect historical periods. Post-OCR correction degrades rather than improves performance. Our methodology provides digital humanities practitioners with guidelines for model selection and quality assessment in historical corpus digitization.

📄 PDF Abstract BibTeX arXiv:2510.06743

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Post-OCR Text Correction for Bulgarian Historical Documents

2024-08-31 · Angel Beshirov, Milena Dobreva, Dimitar Dimitrov, Momchil Hardalov 외

The digitization of historical documents is crucial for preserving the cultural heritage of the society. An important step in this process is converting scanned images to text using Optical Character Recognition (OCR), w…

Optical Character RecognitionOptical Character Recognition (OCR)

Position: On the Methodological Pitfalls of Evaluating Base LLMs for Reasoning

2025-11-13 · Jason Chan, Zhixue Zhao, Robert Gaizauskas arxiv

Existing work investigates the reasoning capabilities of large language models (LLMs) to uncover their limitations, human-like biases and underlying processes. Such studies include evaluations of base LLMs (pre-trained o…

Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents

2025-04-01 · Gavin Greif, Niclas Griesshaber, Robin Greif

We explore how multimodal Large Language Models (mLLMs) can help researchers transcribe historical documents, extract relevant historical information, and construct datasets from historical sources. Specifically, we inve…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+2

Agreement Between Large Language Models, Human Reviewers, and Authors in Evaluating STROBE Checklists for Observational Studies in Rheumatology

2026-03-12 · Emre Bilgin, Ebru Ozturk, Meera Shah, Lisa Traboco 외 arxiv

Introduction: Evaluating compliance with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement can be time-consuming and subjective. This study compares STROBE assessments from large…

Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers

2024-04-04 · YuAn Wang, Xuyang Wu, Hsin-Tai Wu, Zhiqiang Tao 외

The integration of Large Language Models (LLMs) in information retrieval has raised a critical reevaluation of fairness in the text-ranking models. LLMs, such as GPT models and Llama2, have shown effectiveness in natural…

FairnessInformation RetrievalNatural Language UnderstandingRetrieval