Training \& Quality Assessment of an Optical Character Recognition Model for Northern Haida
We are presenting our work on the creation of the first optical character recognition (OCR) model for Northern Haida, also known as Masset or Xaad Kil, a nearly extinct First Nations language spoken in the Haida Gwaii archipelago in British Columbia, Canada. We are addressing the challenges of training an OCR model for a language with an extensive, non-standard Latin character set as follows: (1) We have compared various training approaches and present the results of practical analyses to maximize recognition accuracy and minimize manual labor. An approach using just one or two pages of Source Images directly performed better than the Image Generation approach, and better than models based on three or more pages. Analyses also suggest that a character{'}s frequency is directly correlated with its recognition accuracy. (2) We present an overview of current OCR accuracy analysis tools available. (3) We have ported the once de-facto standardized OCR accuracy tools to be able to cope with Unicode input. Our work adds to a growing body of research on OCR for particularly challenging character sets, and contributes to creating the largest electronic corpus for this severely endangered language.
Code (0)
등록된 구현이 없습니다.
Tasks
Image GenerationOptical Character RecognitionOptical Character Recognition (OCR)Similar Papers 제목 키워드 기반
Optical character recognition quality affects perceived usefulness of historical newspaper clippings
Introduction. We study effect of different quality optical character recognition in interactive information retrieval with a collection of one digitized historical Finnish newspaper. Method. This study is based on the si…
ArticlesInformation RetrievalOptical Character RecognitionOptical Character Recognition (OCR)+1OCR quality affects perceived usefulness of historical newspaper clippings -- a user study
Effects of Optical Character Recognition (OCR) quality on historical information retrieval have so far been studied in data-oriented scenarios regarding the effectiveness of retrieval results. Such studies have either fo…
ArticlesInformation RetrievalOptical Character RecognitionOptical Character Recognition (OCR)+1An Assessment of the Impact of OCR Noise on Language Models
Neural language models are the backbone of modern-day natural language processing applications. Their use on textual heritage collections which have undergone Optical Character Recognition (OCR) is therefore also increas…
Language ModellingOptical Character RecognitionOptical Character Recognition (OCR)Enhancement of text recognition for hanja handwritten documents of Ancient Korea
We implemented a high-performance optical character recognition model for classical handwritten documents using data augmentation with highly variable cropping within the document region. Optical character recognition in…
Data Augmentationobject-detectionObject DetectionOptical Character Recognition+1A Gaussian Process Upsampling Model for Improvements in Optical Character Recognition
Optical Character Recognition and extraction is a key tool in the automatic evaluation of documents in a financial context. However, the image data provided to automated systems can have unreliable quality, and can be in…
Optical Character RecognitionOptical Character Recognition (OCR)