paper-with-me

홈 › Papers

TMIXT: A process flow for Transcribing MIXed handwritten and machine-printed Text

2019-04-28 · Fady Medhat, Mahnaz Mohammadi, Sardar Jaf, Chris G. Willcocks, Toby P. Breckon, Peter Matthews, Andrew Stephen McGough, Georgios Theodoropoulos, Boguslaw Obara

Handling large corpuses of documents is of significant importance in many fields, no more so than in the areas of crime investigation and defence, where an organisation may be presented with a large volume of scanned documents which need to be processed in a finite time. However, this problem is exacerbated both by the volume, in terms of scanned documents and the complexity of the pages, which need to be processed. Often containing many different elements, which each need to be processed and understood. Text recognition, which is a primary task of this process, is usually dependent upon the type of text, being either handwritten or machine-printed. Accordingly, the recognition involves prior classification of the text category, before deciding on the recognition method to be applied. This poses a more challenging task if a document contains both handwritten and machine-printed text. In this work, we present a generic process flow for text recognition in scanned documents containing mixed handwritten and machine-printed text without the need to classify text in advance. We realize the proposed process flow using several open-source image processing and text recognition packages1. The evaluation is performed using a specially developed variant, presented in this work, of the IAM handwriting database, where we achieve an average transcription accuracy of nearly 80% for pages containing both printed and handwritten text.

📄 PDF Abstract BibTeX arXiv:1904.12387

Code (1)

fadymedhat/TMIXT 공식 구현

Similar Papers 제목 키워드 기반

Combining Human and Machine Transcriptions on the Zooniverse Platform

2018-11-01 · WS 2018 11 · Daniel Hanson, Andrea Simenstad

Transcribing handwritten documents to create fully searchable texts is an essential part of the archival process. Traditional text recognition methods, such as optical character recognition (OCR), do not work on handwrit…

Optical Character RecognitionOptical Character Recognition (OCR)

AT-ST: Self-Training Adaptation Strategy for OCR in Domains with Limited Transcriptions

2021-04-27 · Martin Kišš, Karel Beneš, Michal Hradiš

This paper addresses text recognition for domains with limited manual annotations by a simple self-training strategy. Our approach should reduce human annotation effort when target domain data is plentiful, such as when …

Optical Character Recognition (OCR)

An Evaluation of GPT-4V for Transcribing the Urban Renewal Hand-Written Collection

2024-09-11 · Myeong Lee, Julia H. P. Hsu

Between 1960 and 1980, urban renewal transformed many cities, creating vast handwritten records. These documents posed a significant challenge for researchers due to their volume and handwritten nature. The launch of GPT…

Paired Image to Image Translation for Strikethrough Removal From Handwritten Words

2022-01-24 · Raphaela Heil, Ekta Vats, Anders Hast

Transcribing struck-through, handwritten words, for example for the purpose of genetic criticism, can pose a challenge to both humans and machines, due to the obstructive properties of the superimposed strokes. This pape…

Image-to-Image TranslationTranslation

A tailored Handwritten-Text-Recognition System for Medieval Latin

2023-08-18 · Philipp Koch, Gilary Vera Nuñez, Esteban Garces Arias, Christian Heumann 외

The Bavarian Academy of Sciences and Humanities aims to digitize its Medieval Latin Dictionary. This dictionary entails record cards referring to lemmas in medieval Latin, a low-resource language. A crucial step of the d…

Data AugmentationDecoderHandwritten Text RecognitionHTR+2