paper-with-me

홈 › Papers

A Language Modelling Approach to Quality Assessment of OCR’ed Historical Text

2022-06-01 · LREC 2022 6 · Callum Booth, Robert Shoemaker, Robert Gaizauskas

We hypothesise and evaluate a language model-based approach for scoring the quality of OCR transcriptions in the British Library Newspapers (BLN) corpus parts 1 and 2, to identify the best quality OCR for use in further natural language processing tasks, with a wider view to link individual newspaper reports of crime in nineteenth-century London to the Digital Panopticon—a structured repository of criminal lives. We mitigate the absence of gold standard transcriptions of the BLN corpus by utilising a corpus of genre-adjacent texts that capture the common and legal parlance of nineteenth-century London—the Proceedings of the Old Bailey Online—with a view to rank the BLN transcriptions by their OCR quality.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Using SMT for OCR Error Correction of Historical Texts

2016-05-01 · LREC 2016 5 · Haithem Afli, Zhengwei Qiu, Andy Way, P{\'a}raic Sheridan

A trend to digitize historical paper-based archives has emerged in recent years, with the advent of digital optical scanners. A lot of paper-based books, textbooks, magazines, articles, and documents are being transforme…

ArticlesLanguage ModellingMachine TranslationOptical Character Recognition+2

Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality

2025-05-26 · Mohamed Amine Kerkouri, Marouane Tliba, Aladine Chetouani, Nour Aburaed 외

This position paper argues that Mean Opinion Score (MOS), while historically foundational, is no longer sufficient as the sole supervisory signal for multimedia quality assessment models. MOS reduces rich, context-sensit…

An Assessment of the Impact of OCR Noise on Language Models

2022-01-26 · Konstantin Todorov, Giovanni Colavizza

Neural language models are the backbone of modern-day natural language processing applications. Their use on textual heritage collections which have undergone Optical Character Recognition (OCR) is therefore also increas…

Language ModellingOptical Character RecognitionOptical Character Recognition (OCR)

On Modelling Label Uncertainty in Deep Neural Networks: Automatic Estimation of Intra-observer Variability in 2D Echocardiography Quality Assessment

2019-11-02 · Zhibin Liao, Hany Girgis, Amir Abdi, Hooman Vaseli 외

Uncertainty of labels in clinical data resulting from intra-observer variability can have direct impact on the reliability of assessments made by deep neural networks. In this paper, we propose a method for modelling suc…

regression

Evaluating LLMs for Historical Document OCR: A Methodological Framework for Digital Humanities

2025-10-08 · Maria Levchenko arxiv

Digital humanities scholars increasingly use Large Language Models for historical document digitization, yet lack appropriate evaluation frameworks for LLM-based OCR. Traditional metrics fail to capture temporal biases a…