paper-with-me

홈 › Papers

Using SMT for OCR Error Correction of Historical Texts

2016-05-01 · LREC 2016 5 · Haithem Afli, Zhengwei Qiu, Andy Way, P{\'a}raic Sheridan

A trend to digitize historical paper-based archives has emerged in recent years, with the advent of digital optical scanners. A lot of paper-based books, textbooks, magazines, articles, and documents are being transformed into electronic versions that can be manipulated by a computer. For this purpose, Optical Character Recognition (OCR) systems have been developed to transform scanned digital text into editable computer text. However, different kinds of errors in the OCR system output text can be found, but Automatic Error Correction tools can help in performing the quality of electronic texts by cleaning and removing noises. In this paper, we perform a qualitative and quantitative comparison of several error-correction techniques for historical French documents. Experimentation shows that our Machine Translation for Error Correction method is superior to other Language Modelling correction techniques, with nearly 13{\%} relative improvement compared to the initial baseline.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesLanguage ModellingMachine TranslationOptical Character RecognitionOptical Character Recognition (OCR)Translation

Similar Papers 제목 키워드 기반

Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction

2024-07-04 · Laura Manrique-Gómez, Tony Montes, Arturo Rodríguez-Herrera, Rubén Manrique

This paper presents two significant contributions: First, it introduces a novel dataset of 19th-century Latin American newspaper texts, addressing a critical gap in specialized corpora for historical and linguistic analy…

Language ModelingLanguage ModellingLarge Language ModelOptical Character Recognition (OCR)

Toward a Period-Specific Optimized Neural Network for OCR Error Correction of Historical Hebrew Texts

2023-07-30 · Omri Suissa, Maayan Zhitomirsky-Geffet, Avshalom Elmalech

Over the past few decades, large archives of paper-based historical documents, such as books and newspapers, have been digitized using the Optical Character Recognition (OCR) technology. Unfortunately, this broadly used …

Optical Character RecognitionOptical Character Recognition (OCR)

Profiling of OCR'ed Historical Texts Revisited

2017-01-19 · Florian Fink, Klaus-U. Schulz, Uwe Springmann

In the absence of ground truth it is not possible to automatically determine the exact spectrum and occurrences of OCR errors in an OCR'ed text. Yet, for interactive postcorrection of OCR'ed historical printings it is ex…

Optical Character Recognition (OCR)

Optimizing the Neural Network Training for OCR Error Correction of Historical Hebrew Texts

2023-07-30 · Omri Suissa, Avshalom Elmalech, Maayan Zhitomirsky-Geffet

Over the past few decades, large archives of paper-based documents such as books and newspapers have been digitized using Optical Character Recognition. This technology is error-prone, especially for historical documents…

Optical Character RecognitionOptical Character Recognition (OCR)

A Two-Step Approach for Automatic OCR Post-Correction

2020-12-01 · COLING (LaTeCHCLfL, CLFL, LaTeCH) 2020 12 · Robin Schaefer, Clemens Neudecker

The quality of Optical Character Recognition (OCR) is a key factor in the digitisation of historical documents. OCR errors are a major obstacle for downstream tasks and have hindered advances in the usage of the digitise…

Optical Character RecognitionOptical Character Recognition (OCR)Vocal Bursts Valence Prediction