Low-resource Post Processing of Noisy OCR Output for Historical Corpus Digitisation
Code (0)
등록된 구현이 없습니다.
Tasks
Optical Character Recognition (OCR)Similar Papers 제목 키워드 기반
Differentially Private Post-Processing for Fair Regression
This paper describes a differentially private post-processing algorithm for learning fair regressors satisfying statistical parity, addressing privacy concerns of machine learning models trained on sensitive data, as wel…
Density EstimationFairnessregressionDetecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark
This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models ag…
The mapKurator System: A Complete Pipeline for Extracting and Linking Text from Historical Maps
Scanned historical maps in libraries and archives are valuable repositories of geographic data that often do not exist elsewhere. Despite the potential of machine learning tools like the Google Vision APIs for automatica…
Zero-Shot LearningWhen Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts
Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natural Language Processing (NLP) capabilities: low-resource transliterati…
Machine TranslationMultimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents
We explore how multimodal Large Language Models (mLLMs) can help researchers transcribe historical documents, extract relevant historical information, and construct datasets from historical sources. Specifically, we inve…
named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+2