Confidence Prediction for Lexicon-Free OCR
Having a reliable accuracy score is crucial for real world applications of OCR, since such systems are judged by the number of false readings. Lexicon-based OCR systems, which deal with what is essentially a multi-class classification problem, often employ methods explicitly taking into account the lexicon, in order to improve accuracy. However, in lexicon-free scenarios, filtering errors requires an explicit confidence calculation. In this work we show two explicit confidence measurement techniques, and show that they are able to achieve a significant reduction in misreads on both standard benchmarks and a proprietary dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
General ClassificationMulti-class ClassificationOptical Character Recognition (OCR)PredictionSimilar Papers 제목 키워드 기반
LST: Lexicon-Guided Self-Training for Few-Shot Text Classification
Self-training provides an effective means of using an extremely small amount of labeled data to create pseudo-labels for unlabeled data. Many state-of-the-art self-training approaches hinge on different regularization me…
ClassificationFew-Shot Text Classificationtext-classificationText ClassificationData-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
Grapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two…
Self-Supervised Learningtext-to-speechText to SpeechWho Needs Words? Lexicon-Free Speech Recognition
Lexicon-free speech recognition naturally deals with the problem of out-of-vocabulary (OOV) words. In this paper, we show that character-based language models (LM) can perform as well as word-based LMs for speech recogni…
speech-recognitionSpeech RecognitionAnnotation-free Learning of Deep Representations for Word Spotting using Synthetic Data and Self Labeling
Word spotting is a popular tool for supporting the first exploration of historic, handwritten document collections. Today, the best performing methods rely on machine learning techniques, which require a high amount of a…
BIG-bench Machine LearningRetrievalNew Inflectional Lexicons and Training Corpora for Improved Morphosyntactic Annotation of Croatian and Serbian
In this paper we present newly developed inflectional lexcions and manually annotated corpora of Croatian and Serbian. We introduce hrLex and srLex - two freely available inflectional lexicons of Croatian and Serbian - a…
LEMMA