Unsupervised Multi-View Post-OCR Error Correction With Language Models
We investigate post-OCR correction in a setting where we have access to different OCR views of the same document. The goal of this study is to understand if a pretrained language model (LM) can be used in an unsupervised way to reconcile the different OCR views such that their combination contains fewer errors than each individual view. This approach is motivated by scenarios in which unconstrained text generation for error correction is too risky. We evaluated different pretrained LMs on two datasets and found significant gains in realistic scenarios with up to 15% WER improvement over the best OCR view. We also show the importance of domain adaptation for post-OCR correction on out-of-domain documents.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain AdaptationLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)Text GenerationSimilar Papers 제목 키워드 기반
Multi-Input Attention for Unsupervised OCR Correction
We propose a novel approach to OCR post-correction that exploits repeated texts in large corpora both as a source of noisy target outputs for unsupervised training and as a source of evidence when decoding. A sequence-to…
DecoderOptical Character Recognition (OCR)From the Paft to the Fiiture: a Fully Automatic NMT and Word Embeddings Method for OCR Post-Correction
A great deal of historical corpora suffer from errors introduced by the OCR (optical character recognition) methods used in the digitization process. Correcting these errors manually is a time-consuming process and a gre…
BIG-bench Machine LearningMachine TranslationNMTOptical Character Recognition+3An Unsupervised method for OCR Post-Correction and Spelling Normalisation for Finnish
Historical corpora are known to contain errors introduced by OCR (optical character recognition) methods used in the digitization process, often said to be degrading the performance of NLP systems. Correcting these error…
Machine TranslationNMTOptical Character RecognitionOptical Character Recognition (OCR)+1A BERT-based Unsupervised Grammatical Error Correction Framework
Grammatical error correction (GEC) is a challenging task of natural language processing techniques. While more attempts are being made in this approach for universal languages like English or Chinese, relatively little w…
Grammatical Error CorrectionLanguage ModelingLanguage ModellingMulti-class Classification+1Unsupervised domain adaptation for speech recognition with unsupervised error correction
The transcription quality of automatic speech recognition (ASR) systems degrades significantly when transcribing audios coming from unseen domains. We propose an unsupervised error correction method for unsupervised ASR …
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDomain Adaptation+3