paper-with-me

홈 › Papers

CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition

2025-09-24 · Sina J. Semnani, Han Zhang, Xinyan He, Merve Tekgürler, Monica S. Lam arxiv

Accurate text recognition for historical documents can greatly advance the study and preservation of cultural heritage. Existing vision-language models (VLMs), however, are designed for modern, standardized texts and are not equipped to read the diverse languages and scripts, irregular layouts, and frequent degradation found in historical materials. This paper presents CHURRO, a 3B-parameter open-weight VLM specialized for historical text recognition. The model is trained on CHURRO-DS, the largest historical text recognition dataset to date. CHURRO-DS unifies 155 historical corpora comprising 99,491 pages, spanning 22 centuries of textual heritage across 46 language clusters, including historical variants and dead languages. We evaluate several open-weight and closed VLMs and optical character recognition (OCR) systems on CHURRO-DS and find that CHURRO outperforms all other VLMs. On the CHURRO-DS test set, CHURRO achieves 82.3% (printed) and 70.1% (handwritten) normalized Levenshtein similarity, surpassing the second-best model, Gemini 2.5 Pro, by 1.4% and 6.5%, respectively, while being 15.5 times more cost-effective. By releasing the model and dataset, we aim to enable community-driven research to improve the readability of historical texts and accelerate scholarship.

📄 PDF Abstract BibTeX arXiv:2509.19768

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Etymological Wordnet: Tracing The History of Words

2014-05-01 · LREC 2014 5 · Gerard de Melo

Research on the history of words has led to remarkable insights about language and also about the history of human civilization more generally. This paper presents the Etymological Wordnet, the first database that aims a…

Making History Readable

2024-11-26 · Bipasha Banerjee, Jennifer Goyne, William A. Ingram

The Virginia Tech University Libraries (VTUL) Digital Library Platform (DLP) hosts digital collections that offer our users access to a wide variety of documents of historical and cultural importance. These collections a…

Handwriting RecognitionNavigate

Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments

2023-12-11 · Anthony Cintron Roman, Jennifer Wortman Vaughan, Valerie See, Steph Ballard 외

This paper introduces a no-code, machine-readable documentation framework for open datasets, with a focus on responsible AI (RAI) considerations. The framework aims to improve comprehensibility, and usability of open dat…

Decision Making

Context Dependent Semantic Parsing: A Survey

2020-11-02 · COLING 2020 8 · Zhuang Li, Lizhen Qu, Gholamreza Haffari

Semantic parsing is the task of translating natural language utterances into machine-readable meaning representations. Currently, most semantic parsing methods are not able to utilize contextual information (e.g. dialogu…

Semantic ParsingSurvey

After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind

2026-09-15 · Yunpeng Xiong, Ting Zhang arxiv

AI agents increasingly act through agent skills, i.e., natural-language instructions, that direct a host agent toward shell, network, credential, file, and process actions, and public registries distribute them at scale.…