paper-with-me

홈 › Papers

Page Layout Analysis System for Unconstrained Historic Documents

2021-02-23 · Oldřich Kodym, Michal Hradiš

Extraction of text regions and individual text lines from historic documents is necessary for automatic transcription. We propose extending a CNN-based text baseline detection system by adding line height and text block boundary predictions to the model output, allowing the system to extract more comprehensive layout information. We also show that pixel-wise text orientation prediction can be used for processing documents with multiple text orientations. We demonstrate that the proposed method performs well on the cBAD baseline detection dataset. Additionally, we benchmark the method on newly introduced PERO layout dataset which we also make public.

📄 PDF Abstract BibTeX arXiv:2102.11838

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FP-THD: Full page transcription of historical documents

2026-01-20 · H Neji, J Nogueras-Iso, J Lacasta, MÁ Latre 외 arxiv

The transcription of historical documents written in Latin in XV and XVI centuries has special challenges as it must maintain the characters and special symbols that have distinct meanings to ensure that historical texts…

Page Layout Analysis of Text-heavy Historical Documents: a Comparison of Textual and Visual Approaches

2022-12-12 · Najem-Meyer Sven, Romanello Matteo

Page layout analysis is a fundamental step in document processing which enables to segment a page into regions of interest. With highly complex layouts and mixed scripts, scholarly commentaries are text-heavy documents w…

Position

PereStruct: Multimodal Semantic Assembly for Robust Historical Document Parsing

2026-06-03 · Maksim Shandybo, Ivan Bespalov, Daniil Yefimov, Marina Kosheleva 외 arxiv

Parsing historical documents with complex, non-standard layouts remains a fundamental bottleneck in large-scale archival digitization. Unlike modern typography, historical newspapers exhibit severe physical degradation a…

Semantic Similarity

Page image classification for content-specific data processing

2025-07-11 · Kateryna Lutsai arxiv

Digitization projects in humanities often generate vast quantities of page images from historical documents, presenting significant challenges for manual sorting and analysis. These archives contain diverse content, incl…

Image Classification

Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts

2026-08-19 · Kartik Chincholikar, Kaushik Gopalan, Mihir Hasabnis arxiv

Digitizing the text from handwritten historical manuscripts is required to make them easily accessible, preservable, and to enable historical scholars to study them in new ways. Historical manuscripts, however, often exh…