paper-with-me

Papers

Multi-modular domain-tailored OCR post-correction

2017-09-01 · EMNLP 2017 9 · Sarah Schulz, Jonas Kuhn

One of the main obstacles for many Digital Humanities projects is the low data availability. Texts have to be digitized in an expensive and time consuming process whereas Optical Character Recognition (OCR) post-correction is one of the time-critical factors. At the example of OCR post-correction, we show the adaptation of a generic system to solve a specific problem with little data. The system accounts for a diversity of errors encountered in OCRed texts coming from different time periods in the domain of literature. We show that the combination of different approaches, such as e.g. Statistical Machine Translation and spell checking, with the help of a ranking mechanism tremendously improves over single-handed approaches. Since we consider the accessibility of the resulting tool as a crucial part of Digital Humanities collaborations, we describe the workflow we suggest for efficient text recognition and subsequent automatic and manual post-correction

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityMachine TranslationOptical Character RecognitionOptical Character Recognition (OCR)Translation

Similar Papers 제목 키워드 기반

Speech Recognition on TV Series with Video-guided Post-Correction

2025-06-08 · Haoyuan Yang, Yue Zhang, Liqiang Jing

Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assistive technologies. However, ASR systems…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Large Language Modelspeech-recognition+1

Improving Diffusion Posterior Samplers with Lagged Temporal Corrections for Image Restoration

2026-05-12 · Davide Evangelista, Elena Morotti, Francesco Pivi, Maurizio Gabbrielli arxiv

Diffusion-based posterior sampling (PS) is a leading framework for imaging inverse problems, combining learned priors with measurement constraints. Yet, its standard formulations rely on instantaneous data-consistent est…

Image Restoration

Unsupervised Multi-View Post-OCR Error Correction With Language Models

2021-11-01 · EMNLP 2021 11 · Harsh Gupta, Luciano del Corro, Samuel Broscheit, Johannes Hoffart 외

We investigate post-OCR correction in a setting where we have access to different OCR views of the same document. The goal of this study is to understand if a pretrained language model (LM) can be used in an unsupervised…

Domain AdaptationLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)+1

CFD-copilot: leveraging domain-adapted large language model and model context protocol to enhance simulation automation

2025-12-08 · Zhehao Dong, Shanghai Du, Zhen Lu, Yue Yang arxiv

Configuring computational fluid dynamics (CFD) simulations requires significant expertise in physics modeling and numerical methods, posing a barrier to non-specialists. Although automating scientific tasks with large la…

EXCGEC: A Benchmark of Edit-wise Explainable Chinese Grammatical Error Correction

2024-07-01 · Jingheng Ye, Shang Qin, Yinghui Li, Xuxin Cheng 외

Existing studies explore the explainability of Grammatical Error Correction (GEC) in a limited scenario, where they ignore the interaction between corrections and explanations. To bridge the gap, this paper introduces th…

Grammatical Error Correction