paper-with-me

Papers

Low-resource Post Processing of Noisy OCR Output for Historical Corpus Digitisation

2018-05-01 · LREC 2018 5 · Caitlin Richter, Matthew Wickes, Deniz Beser, Mitch Marcus
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Optical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Differentially Private Post-Processing for Fair Regression

2024-05-07 · Ruicheng Xian, Qiaobo Li, Gautam Kamath, Han Zhao

This paper describes a differentially private post-processing algorithm for learning fair regressors satisfying statistical parity, addressing privacy concerns of machine learning models trained on sensitive data, as wel…

Density EstimationFairnessregression

Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark

2025-10-22 · Yu Wu, Ke Shu, Jonas Fischer, Lidia Pivovarova 외 arxiv

This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models ag…

The mapKurator System: A Complete Pipeline for Extracting and Linking Text from Historical Maps

2023-06-29 · Jina Kim, Zekun Li, Yijun Lin, Min Namgung 외

Scanned historical maps in libraries and archives are valuable repositories of geographic data that often do not exist elsewhere. Despite the potential of machine learning tools like the Google Vision APIs for automatica…

Zero-Shot Learning

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

2026-07-04 · Nguyen Kim Hai Bui, Md. Easin Arafat, Tamás Gábor Orosz, Mufti Mahmud arxiv

Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natural Language Processing (NLP) capabilities: low-resource transliterati…

Machine Translation

Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents

2025-04-01 · Gavin Greif, Niclas Griesshaber, Robin Greif

We explore how multimodal Large Language Models (mLLMs) can help researchers transcribe historical documents, extract relevant historical information, and construct datasets from historical sources. Specifically, we inve…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+2