paper-with-me

Papers

Unsupervised Multi-View Post-OCR Error Correction With Language Models

2021-11-01 · EMNLP 2021 11 · Harsh Gupta, Luciano del Corro, Samuel Broscheit, Johannes Hoffart, Eliot Brenner

We investigate post-OCR correction in a setting where we have access to different OCR views of the same document. The goal of this study is to understand if a pretrained language model (LM) can be used in an unsupervised way to reconcile the different OCR views such that their combination contains fewer errors than each individual view. This approach is motivated by scenarios in which unconstrained text generation for error correction is too risky. We evaluated different pretrained LMs on two datasets and found significant gains in realistic scenarios with up to 15% WER improvement over the best OCR view. We also show the importance of domain adaptation for post-OCR correction on out-of-domain documents.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)Text Generation

Similar Papers 제목 키워드 기반

Multi-Input Attention for Unsupervised OCR Correction

2018-07-01 · ACL 2018 7 · Rui Dong, David Smith

We propose a novel approach to OCR post-correction that exploits repeated texts in large corpora both as a source of noisy target outputs for unsupervised training and as a source of evidence when decoding. A sequence-to…

DecoderOptical Character Recognition (OCR)

From the Paft to the Fiiture: a Fully Automatic NMT and Word Embeddings Method for OCR Post-Correction

2019-10-12 · RANLP 2019 9 · Mika Hämäläinen, Simon Hengchen

A great deal of historical corpora suffer from errors introduced by the OCR (optical character recognition) methods used in the digitization process. Correcting these errors manually is a time-consuming process and a gre…

BIG-bench Machine LearningMachine TranslationNMTOptical Character Recognition+3

An Unsupervised method for OCR Post-Correction and Spelling Normalisation for Finnish

2020-11-06 · NoDaLiDa 2021 5 · Quan Duong, Mika Hämäläinen, Simon Hengchen

Historical corpora are known to contain errors introduced by OCR (optical character recognition) methods used in the digitization process, often said to be degrading the performance of NLP systems. Correcting these error…

Machine TranslationNMTOptical Character RecognitionOptical Character Recognition (OCR)+1

A BERT-based Unsupervised Grammatical Error Correction Framework

2023-03-30 · Nankai Lin, Hongbin Zhang, Menglan Shen, Yu Wang 외

Grammatical error correction (GEC) is a challenging task of natural language processing techniques. While more attempts are being made in this approach for universal languages like English or Chinese, relatively little w…

Grammatical Error CorrectionLanguage ModelingLanguage ModellingMulti-class Classification+1

Unsupervised domain adaptation for speech recognition with unsupervised error correction

2022-09-24 · Long Mai, Julie Carson-Berndsen

The transcription quality of automatic speech recognition (ASR) systems degrades significantly when transcribing audios coming from unseen domains. We propose an unsupervised error correction method for unsupervised ASR …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDomain Adaptation+3