paper-with-me

홈 › Papers

State of the Art Optical Character Recognition of 19th Century Fraktur Scripts using Open Source Engines

2018-10-08 · Christian Reul, Uwe Springmann, Christoph Wick, Frank Puppe

In this paper we evaluate Optical Character Recognition (OCR) of 19th century Fraktur scripts without book-specific training using mixed models, i.e. models trained to recognize a variety of fonts and typesets from previously unseen sources. We describe the training process leading to strong mixed OCR models and compare them to freely available models of the popular open source engines OCRopus and Tesseract as well as the commercial state of the art system ABBYY. For evaluation, we use a varied collection of unseen data from books, journals, and a dictionary from the 19th century. The experiments show that training mixed models with real data is superior to training with synthetic data and that the novel OCR engine Calamari outperforms the other engines considerably, on average reducing ABBYYs character error rate (CER) by over 70%, resulting in an average CER below 1%.

📄 PDF Abstract BibTeX arXiv:1810.03436

Code (1)

chreul/19th-century-fraktur-OCR 공식 구현

Tasks

Optical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin

2018-09-14 · Uwe Springmann, Christian Reul, Stefanie Dipper, Johannes Baiter

In this paper we describe a dataset of German and Latin \textit{ground truth} (GT) for historical OCR in the form of printed text line images paired with their transcription. This dataset, called \textit{GT4HistOCR}, con…

Optical Character Recognition (OCR)

Improving Optical Character Recognition of Finnish Historical Newspapers with a Combination of Fraktur \& Antiqua Models and Image Preprocessing

2017-05-01 · WS 2017 5 · Mika Koistinen, Kimmo Kettunen, Tuula P{\"a}{\"a}kk{\"o}nen
Boundary DetectionInformation RetrievalMachine TranslationNamed Entity Recognition (NER)+2

Mixed Model OCR Training on Historical Latin Script for Out-of-the-Box Recognition and Finetuning

2021-06-15 · Christian Reul, Christoph Wick, Maximilian Nöth, Andreas Büttner 외

In order to apply Optical Character Recognition (OCR) to historical printings of Latin script fully automatically, we report on our efforts to construct a widely-applicable polyfont recognition model yielding text with a…

Data AugmentationOptical Character RecognitionOptical Character Recognition (OCR)

Calamari - A High-Performance Tensorflow-based Deep Learning Package for Optical Character Recognition

2018-07-05 · Christoph Wick, Christian Reul, Frank Puppe

Optical Character Recognition (OCR) on contemporary and historical data is still in the focus of many researchers. Especially historical prints require book specific trained OCR models to achieve applicable results (Spri…

GPUOptical Character RecognitionOptical Character Recognition (OCR)

Optical Character Recognition of 19th Century Classical Commentaries: the Current State of Affairs

2021-10-13 · Matteo Romanello, Sven Najem-Meyer, Bruce Robertson

Together with critical editions and translations, commentaries are one of the main genres of publication in literary and textual scholarship, and have a century-long tradition. Yet, the exploitation of thousands of digit…

Optical Character RecognitionOptical Character Recognition (OCR)