paper-with-me

홈 › Papers

Training Kindai OCR with parallel textline images and self-attention feature distance-based loss

2025-08-12 · Anh Le, Asanobu Kitamoto arxiv

Kindai documents, written in modern Japanese from the late 19th to early 20th century, hold significant historical value for researchers studying societal structures, daily life, and environmental conditions of that period. However, transcribing these documents remains a labor-intensive and time-consuming task, resulting in limited annotated data for training optical character recognition (OCR) systems. This research addresses this challenge of data scarcity by leveraging parallel textline images - pairs of original Kindai text and their counterparts in contemporary Japanese fonts - to augment training datasets. We introduce a distance-based objective function that minimizes the gap between self-attention features of the parallel image pairs. Specifically, we explore Euclidean distance and Maximum Mean Discrepancy (MMD) as domain adaptation metrics. Experimental results demonstrate that our method reduces the character error rate (CER) by 2.23% and 3.94% over a Transformer-based OCR baseline when using Euclidean distance and MMD, respectively. Furthermore, our approach improves the discriminative quality of self-attention representations, leading to more effective OCR performance for historical documents.

📄 PDF Abstract BibTeX arXiv:2508.08537

Code (0)

등록된 구현이 없습니다.

Tasks

Domain Adaptation

Similar Papers 제목 키워드 기반

Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document Understanding

2022-12-19 · Haoli Bai, Zhiguang Liu, Xiaojun Meng, Wentao Li 외

Unsupervised pre-training on millions of digital-born or scanned documents has shown promising advances in visual document understanding~(VDU). While various vision-language pre-training objectives are studied in existin…

Contrastive Learningdocument understandingOptical Character Recognition (OCR)Unsupervised Pre-training

ScanSSD: Scanning Single Shot Detector for Mathematical Formulas in PDF Document Images

2020-03-18 · Parag Mali, Puneeth Kukkadapu, Mahshad Mahdavi, Richard Zanibbi

We introduce the Scanning Single Shot Detector (ScanSSD) for locating math formulas offset from text and embedded in textlines. ScanSSD uses only visual features for detection: no formatting or typesetting information su…

Math

Geometric Representation Learning for Document Image Rectification

2022-10-15 · Hao Feng, Wengang Zhou, Jiajun Deng, Yuechen Wang 외

In document image rectification, there exist rich geometric constraints between the distorted image and the ground truth one. However, such geometric constraints are largely ignored in existing advanced solutions, which …

Representation Learning

ICDAR 2019 Historical Document Reading Challenge on Large Structured Chinese Family Records

2019-03-08 · Rajkumar Saini, Derek Dobson, Jon Morrey, Marcus Liwicki 외

We propose a Historical Document Reading Challenge on Large Chinese Structured Family Records, in short ICDAR2019 HDRC CHINESE. The objective of the proposed competition is to recognize and analyze the layout, and finall…

Few Shots Are All You Need: A Progressive Few Shot Learning Approach for Low Resource Handwritten Text Recognition

2021-07-21 · Mohamed Ali Souibgui, Alicia Fornés, Yousri Kessentini, Beáta Megyesi

Handwritten text recognition in low resource scenarios, such as manuscripts with rare alphabets, is a challenging problem. The main difficulty comes from the very few annotated data and the limited linguistic information…

AllFew-Shot LearningHandwriting RecognitionHandwritten Text Recognition