paper-with-me

Papers

Historical Document Processing: Historical Document Processing: A Survey of Techniques, Tools, and Trends

2020-02-15 · James P. Philips, Nasseh Tabrizi

Historical Document Processing is the process of digitizing written material from the past for future use by historians and other scholars. It incorporates algorithms and software tools from various subfields of computer science, including computer vision, document analysis and recognition, natural language processing, and machine learning, to convert images of ancient manuscripts, letters, diaries, and early printed texts automatically into a digital format usable in data mining and information retrieval systems. Within the past twenty years, as libraries, museums, and other cultural heritage institutions have scanned an increasing volume of their historical document archives, the need to transcribe the full text from these collections has become acute. Since Historical Document Processing encompasses multiple sub-domains of computer science, knowledge relevant to its purpose is scattered across numerous journals and conference proceedings. This paper surveys the major phases of, standard algorithms, tools, and datasets in the field of Historical Document Processing, discusses the results of a literature review, and finally suggests directions for further research.

📄 PDF Abstract BibTeX arXiv:2002.06300

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

Predicting the Original Appearance of Damaged Historical Documents

2024-12-16 · Zhenhua Yang, Dezhi Peng, Yongxin Shi, Yuyi Zhang 외

Historical documents encompass a wealth of cultural treasures but suffer from severe damages including character missing, paper damage, and ink erosion over time. However, existing document processing methods primarily f…

Binarization

Language Resources for Historical Newspapers: the Impresso Collection

2020-05-01 · LREC 2020 5 · Maud Ehrmann, Matteo Romanello, Simon Clematide, Phillip Benjamin Str{\"o}bel 외

Following decades of massive digitization, an unprecedented amount of historical document facsimiles can now be retrieved and accessed via cultural heritage online portals. If this represents a huge step forward in terms…

Large Language Models for Summarizing Czech Historical Documents and Beyond

2025-08-14 · Václav Tran, Jakub Šmíd, Jiří Martínek, Ladislav Lenc 외 arxiv

Text summarization is the task of shortening a larger body of text into a concise version while retaining its essential meaning and key information. While summarization has been significantly explored in English and othe…

Text Summarization

HERITAGE: An End-to-End Web Platform for Processing Korean Historical Documents in Hanja

2025-01-21 · Seyoung Song, Haneul Yoo, Jiho Jin, Kyunghyun Cho 외

While Korean historical documents are invaluable cultural heritage, understanding those documents requires in-depth Hanja expertise. Hanja is an ancient language used in Korea before the 20th century, whose characters we…

document understandingMachine Translationnamed-entity-recognitionNamed Entity Recognition+1

FP-THD: Full page transcription of historical documents

2026-01-20 · H Neji, J Nogueras-Iso, J Lacasta, MÁ Latre 외 arxiv

The transcription of historical documents written in Latin in XV and XVI centuries has special challenges as it must maintain the characters and special symbols that have distinct meanings to ensure that historical texts…