paper-with-me

홈 › Papers

Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing

2026-05-25 · Kateryna Lutsai, Dana Křivánková, Pavel Straňák, David Novák arxiv

Purpose: Digitization projects in the humanities produce vast, heterogeneous archives of historical documents, making manual sorting impractical at scale. This work addresses the need for an automated system to classify scanned page images based on visual content type - text, tables, and graphics - enabling content-specific downstream processing such as Optical Character Recognition (OCR) or structured data extraction. Methods: An image classification system was developed and evaluated on a dataset of over 48,000 annotated historical page images from century-old Czech archaeological archives, refined through four successive annotation stages with domain-expert review. A Random Forest Classifier baseline was established using hand-crafted image features. Subsequently, deep learning architectures were fine-tuned and compared: Convolutional Neural Networks (EfficientNetV2, RegNetY), Vision and Document Image Transformers (ViT, DiT), and multimodal CLIP models. An 11-category label scheme was designed collaboratively with domain experts and evaluated via five-fold cross-validation. Results: The feature-based baseline achieved approximately 75% accuracy. Fine-tuned CNNs and Transformers substantially outperformed it, with RegNetY-16GF achieving 99.16% and ViT-large 99.12% Top-1 accuracy on the held-out test set. CLIP ViT-B/16 reached 99.14% with optimized text descriptions. Conclusion: Image-only models, particularly RegNetY-16GF, deliver near-perfect classification accuracy and produce consistent labels across 649,508 unlabeled archival pages with over 90% inter-model agreement. Fine-tuned CLIP, despite competitive test-set accuracy, showed under 65% agreement with image-only models on unlabeled data, making it less suitable for deployment. The final models, annotated dataset, and software are publicly available under open-source licenses.

📄 PDF Abstract BibTeX arXiv:2606.07558

Code (0)

등록된 구현이 없습니다.

Tasks

Image Classification

Similar Papers 제목 키워드 기반

Sequence-aware multimodal page classification of Brazilian legal documents

2022-07-02 · Pedro H. Luz de Araujo, Ana Paula G. S. de Almeida, Fabricio A. Braz, Nilton C. da Silva 외

The Brazilian Supreme Court receives tens of thousands of cases each semester. Court employees spend thousands of hours to execute the initial analysis and classification of those cases -- which takes effort away from po…

ClassificationManagementOptical Character RecognitionOptical Character Recognition (OCR)

Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models

2025-02-18 · Jonathan Bourne

Oscar Wilde said, "The difference between literature and journalism is that journalism is unreadable, and literature is not read." Unfortunately, The digitally archived journalism of Oscar Wilde's 19th century often has …

Image to textOptical Character RecognitionOptical Character Recognition (OCR)Topic Classification

EDDA-Coordinata: An Annotated Dataset of Historical Geographic Coordinates

2026-02-27 · Ludovic Moncla, Pierre Nugues, Thierry Joliveau, Katherine McDonough arxiv

This paper introduces a dataset of enriched geographic coordinates retrieved from Diderot and d'Alembert's eighteenth-century Encyclopedie. Automatically recovering geographic coordinates from historical texts is a compl…

Seventeenth-Century Spanish American Notary Records for Fine-Tuning Spanish Large Language Models

2024-06-09 · Shraboni Sarker, Ahmad Tamim Hamad, Hulayyil Alshammari, Viviana Grieco 외

Large language models have gained tremendous popularity in domains such as e-commerce, finance, healthcare, and education. Fine-tuning is a common approach to customize an LLM on a domain-specific dataset for a desired d…

Language ModelingLanguage ModellingMasked Language Modeling

Historical Ink: Semantic Shift Detection for 19th Century Spanish

2024-07-08 · Tony Montes, Laura Manrique-Gómez, Rubén Manrique

This paper explores the evolution of word meanings in 19th-century Spanish texts, with an emphasis on Latin American Spanish, using computational linguistics techniques. It addresses the Semantic Shift Detection (SSD) ta…

Masked Language ModelingSemantic Shift DetectionSemantic Similarity