Multi-modular domain-tailored OCR post-correction
One of the main obstacles for many Digital Humanities projects is the low data availability. Texts have to be digitized in an expensive and time consuming process whereas Optical Character Recognition (OCR) post-correction is one of the time-critical factors. At the example of OCR post-correction, we show the adaptation of a generic system to solve a specific problem with little data. The system accounts for a diversity of errors encountered in OCRed texts coming from different time periods in the domain of literature. We show that the combination of different approaches, such as e.g. Statistical Machine Translation and spell checking, with the help of a ranking mechanism tremendously improves over single-handed approaches. Since we consider the accessibility of the resulting tool as a crucial part of Digital Humanities collaborations, we describe the workflow we suggest for efficient text recognition and subsequent automatic and manual post-correction
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityMachine TranslationOptical Character RecognitionOptical Character Recognition (OCR)TranslationSimilar Papers 제목 키워드 기반
Speech Recognition on TV Series with Video-guided Post-Correction
Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assistive technologies. However, ASR systems…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Large Language Modelspeech-recognition+1Improving Diffusion Posterior Samplers with Lagged Temporal Corrections for Image Restoration
Diffusion-based posterior sampling (PS) is a leading framework for imaging inverse problems, combining learned priors with measurement constraints. Yet, its standard formulations rely on instantaneous data-consistent est…
Image RestorationUnsupervised Multi-View Post-OCR Error Correction With Language Models
We investigate post-OCR correction in a setting where we have access to different OCR views of the same document. The goal of this study is to understand if a pretrained language model (LM) can be used in an unsupervised…
Domain AdaptationLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)+1CFD-copilot: leveraging domain-adapted large language model and model context protocol to enhance simulation automation
Configuring computational fluid dynamics (CFD) simulations requires significant expertise in physics modeling and numerical methods, posing a barrier to non-specialists. Although automating scientific tasks with large la…
EXCGEC: A Benchmark of Edit-wise Explainable Chinese Grammatical Error Correction
Existing studies explore the explainability of Grammatical Error Correction (GEC) in a limited scenario, where they ignore the interaction between corrections and explanations. To bridge the gap, this paper introduces th…
Grammatical Error Correction