Handling Heavily Abbreviated Manuscripts: HTR engines vs text normalisation approaches
Although abbreviations are fairly common in handwritten sources, particularly in medieval and modern Western manuscripts, previous research dealing with computational approaches to their expansion is scarce. Yet abbreviations present particular challenges to computational approaches such as handwritten text recognition and natural language processing tasks. Often, pre-processing ultimately aims to lead from a digitised image of the source to a normalised text, which includes expansion of the abbreviations. We explore different setups to obtain such a normalised text, either directly, by training HTR engines on normalised (i.e., expanded, disabbreviated) text, or by decomposing the process into discrete steps, each making use of specialist models for recognition, word segmentation and normalisation. The case studies considered here are drawn from the medieval Latin tradition.
Code (0)
등록된 구현이 없습니다.
Tasks
Handwritten Text RecognitionHTRSimilar Papers 제목 키워드 기반
Cleansing Jewel: A Neural Spelling Correction Model Built On Google OCR-ed Tibetan Manuscripts
Scholars in the humanities rely heavily on ancient manuscripts to study history, religion, and socio-political structures in the past. Many efforts have been devoted to digitizing these precious manuscripts using OCR tec…
Optical Character RecognitionOptical Character Recognition (OCR)Spelling CorrectionFrom exemplar to copy: the scribal appropriation of a Hadewijch manuscript computationally explored
This study is devoted to two of the oldest known manuscripts in which the oeuvre of the medieval mystical author Hadewijch has been preserved: Brussels, KBR, 2879-2880 (ms. A) and Brussels, KBR, 2877-2878 (ms. B). On the…
ICDAR 2024 Competition on Few-Shot and Many-Shot Layout Segmentation of Ancient Manuscripts (SAM)
Layout analysis is a critical aspect of Document Image Analysis, particularly when it comes to ancient manuscripts. It serves as a foundational step in streamlining subsequent tasks such as optical character recognition …
DiversityDocument Layout AnalysisFew-Shot LearningOptical Character RecognitionTextual analysis of artificial intelligence manuscripts reveals features associated with peer review outcome
We analysed a dataset of scientific manuscripts that were submitted to various conferences in artificial intelligence. We performed a combination of semantic, lexical and psycholinguistic analyses of the full text of the…
Semantic SimilaritySemantic Textual SimilarityAccelerating Text Communication via Abbreviated Sentence Input
Typing every character in a text message may require more time or effort than strictly necessary. Skipping spaces or other characters may be able to speed input and reduce a user{'}s physical input effort. This can be pa…
Sentence