paper-with-me

홈 › Papers

OCR Post-Correction Evaluation of Early Dutch Books Online - Revisited

2016-05-01 · LREC 2016 5 · Martin Reynaert

We present further work on evaluation of the fully automatic post-correction of Early Dutch Books Online, a collection of 10,333 18th century books. In prior work we evaluated the new implementation of Text-Induced Corpus Clean-up (TICCL) on the basis of a single book Gold Standard derived from this collection. In the current paper we revisit the same collection on the basis of a sizeable 1020 item random sample of OCR post-corrected strings from the full collection. Both evaluations have their own stories to tell and lessons to teach.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Synergy of Nederlab and

2014-05-01 · LREC 2014 5 · Martin Reynaert

In two concurrent projects in the Netherlands we are further developing TICCL or Text-Induced Corpus Clean-up. In project Nederlab TICCL is set to work on diachronic Dutch text. To this end it has been equipped with the …

Optical Character Recognition (OCR)

LAREX - A semi-automatic open-source Tool for Layout Analysis and Region Extraction on Early Printed Books

2017-01-20 · Christian Reul, Uwe Springmann, Frank Puppe

A semi-automatic open-source tool for layout analysis on early printed books is presented. LAREX uses a rule based connected components approach which is very fast, easily comprehensible for the user and allows an intuit…

Optical Character Recognition (OCR)

Multi-Input Attention for Unsupervised OCR Correction

2018-07-01 · ACL 2018 7 · Rui Dong, David Smith

We propose a novel approach to OCR post-correction that exploits repeated texts in large corpora both as a source of noisy target outputs for unsupervised training and as a source of evidence when decoding. A sequence-to…

DecoderOptical Character Recognition (OCR)

OCR Post Correction for Endangered Language Texts

2020-11-10 · EMNLP 2020 11 · Shruti Rijhwani, Antonios Anastasopoulos, Graham Neubig

There is little to no data available to build natural language processing models for most endangered languages. However, textual data in these languages often exists in formats that are not machine-readable, such as pape…

Optical Character Recognition (OCR)

A Hardware-Aware, Per-Layer Methodology for Post-Training Quantization of Large Language Models

2026-05-14 · Earl Killian arxiv

Scaled Outer Product (SOP) is a post-training quantization methodology for large language model weights, designed to deliver near-lossless fidelity at 4.5--6 bits per weight on hardware with per-layer LUT decode. The met…