paper-with-me

홈 › Papers

Neural OCR Post-Hoc Correction of Historical Corpora

2021-02-01 · Lijun Lyu, Maria Koutraki, Martin Krickl, Besnik Fetahu

Optical character recognition (OCR) is crucial for a deeper access to historical collections. OCR needs to account for orthographic variations, typefaces, or language evolution (i.e., new letters, word spellings), as the main source of character, word, or word segmentation transcription errors. For digital corpora of historical prints, the errors are further exacerbated due to low scan quality and lack of language standardization. For the task of OCR post-hoc correction, we propose a neural approach based on a combination of recurrent (RNN) and deep convolutional network (ConvNet) to correct OCR transcription errors. At character level we flexibly capture errors, and decode the corrected output based on a novel attention mechanism. Accounting for the input and output similarity, we propose a new loss function that rewards the model's correcting behavior. Evaluation on a historical book corpus in German language shows that our models are robust in capturing diverse OCR transcription errors and reduce the word error rate of 32.3% by more than 89%.

📄 PDF Abstract BibTeX arXiv:2102.00583

Code (1)

GarfieldLyu/OCR_POST_DE 공식 구현 pytorch

Tasks

Optical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Toward a Period-Specific Optimized Neural Network for OCR Error Correction of Historical Hebrew Texts

2023-07-30 · Omri Suissa, Maayan Zhitomirsky-Geffet, Avshalom Elmalech

Over the past few decades, large archives of paper-based historical documents, such as books and newspapers, have been digitized using the Optical Character Recognition (OCR) technology. Unfortunately, this broadly used …

Optical Character RecognitionOptical Character Recognition (OCR)

Multi-Input Attention for Unsupervised OCR Correction

2018-07-01 · ACL 2018 7 · Rui Dong, David Smith

We propose a novel approach to OCR post-correction that exploits repeated texts in large corpora both as a source of noisy target outputs for unsupervised training and as a source of evidence when decoding. A sequence-to…

DecoderOptical Character Recognition (OCR)

Optimizing the Neural Network Training for OCR Error Correction of Historical Hebrew Texts

2023-07-30 · Omri Suissa, Avshalom Elmalech, Maayan Zhitomirsky-Geffet

Over the past few decades, large archives of paper-based documents such as books and newspapers have been digitized using Optical Character Recognition. This technology is error-prone, especially for historical documents…

Optical Character RecognitionOptical Character Recognition (OCR)

From the Paft to the Fiiture: a Fully Automatic NMT and Word Embeddings Method for OCR Post-Correction

2019-10-12 · RANLP 2019 9 · Mika Hämäläinen, Simon Hengchen

A great deal of historical corpora suffer from errors introduced by the OCR (optical character recognition) methods used in the digitization process. Correcting these errors manually is a time-consuming process and a gre…

BIG-bench Machine LearningMachine TranslationNMTOptical Character Recognition+3

OCR and post-correction of historical Finnish texts

2017-05-01 · WS 2017 5 · Senka Drobac, Pekka Kauppinen, Krister Lind{\'e}n
Optical Character Recognition (OCR)Spelling Correction