paper-with-me

홈 › Papers

HANA: A HAndwritten NAme Database for Offline Handwritten Text Recognition

2021-01-22 · Christian M. Dahl, Torben Johansen, Emil N. Sørensen, Simon Wittrock

Methods for linking individuals across historical data sets, typically in combination with AI based transcription models, are developing rapidly. Probably the single most important identifier for linking is personal names. However, personal names are prone to enumeration and transcription errors and although modern linking methods are designed to handle such challenges, these sources of errors are critical and should be minimized. For this purpose, improved transcription methods and large-scale databases are crucial components. This paper describes and provides documentation for HANA, a newly constructed large-scale database which consists of more than 3.3 million names. The database contain more than 105 thousand unique names with a total of more than 1.1 million images of personal names, which proves useful for transfer learning to other settings. We provide three examples hereof, obtaining significantly improved transcription accuracy on both Danish and US census data. In addition, we present benchmark results for deep learning models automatically transcribing the personal names from the scanned documents. Through making more challenging large-scale databases publicly available we hope to foster more sophisticated, accurate, and robust models for handwritten text recognition.

📄 PDF Abstract BibTeX arXiv:2101.10862

Code (1)

torbensdjohansen/hana 공식 구현 pytorch

Tasks

Handwritten Text RecognitionTransfer Learning

Similar Papers 제목 키워드 기반

Development of Comprehensive Devnagari Numeral and Character Database for Offline Handwritten Character Recognition

2013-08-17 · Vikas J. Dongre, Vijay H. Mankar

In handwritten character recognition, benchmark database plays an important role in evaluating the performance of various algorithms and the results obtained by various researchers. In Devnagari script, there is lack of …

Handwriting Recognition

Handwritten Urdu Character Recognition using 1-Dimensional BLSTM Classifier

2017-05-15 · Saad Bin Ahmed, Saeeda Naz, Salahuddin Swati, Muhammad Imran Razzak

The recognition of cursive script is regarded as a subtle task in optical character recognition due to its varied representation. Every cursive script has different nature and associated challenges. As Urdu is one of cur…

Optical Character RecognitionOptical Character Recognition (OCR)

uTHCD: A New Benchmarking for Tamil Handwritten OCR

2021-03-13 · Noushath Shaffi, Faizal Hajamohideen

Handwritten character recognition is a challenging research in the field of document image analysis over many decades due to numerous reasons such as large writing styles variation, inherent noise in data, expansive appl…

BenchmarkingOptical Character Recognition (OCR)

Boosting offline handwritten text recognition in historical documents with few labeled lines

2020-12-04 · José Carlos Aradillas, Juan José Murillo-Fuentes, Pablo M. Olmos

In this paper, we face the problem of offline handwritten text recognition (HTR) in historical documents when few labeled samples are available and some of them contain errors in the train set. Three main contributions a…

Data AugmentationHandwritten Text RecognitionHTRTransfer Learning

Classification of Handwritten Names of Cities and Handwritten Text Recognition using Various Deep Learning Models

2021-02-09 · Daniyar Nurseitov, Kairat Bostanbekov, Maksat Kanatov, Anel Alimova 외

This article discusses the problem of handwriting recognition in Kazakh and Russian languages. This area is poorly studied since in the literature there are almost no works in this direction. We have tried to describe va…

Handwriting RecognitionHandwritten Text Recognition