Correction of OCR Word Segmentation Errors in Articles from the ACL Collection through Neural Machine Translation Methods
Code (0)
등록된 구현이 없습니다.
Tasks
ArticlesGrammatical Error CorrectionKeyword ExtractionMachine TranslationOptical Character Recognition (OCR)Relation ExtractionTranslationSimilar Papers 제목 키워드 기반
Neural OCR Post-Hoc Correction of Historical Corpora
Optical character recognition (OCR) is crucial for a deeper access to historical collections. OCR needs to account for orthographic variations, typefaces, or language evolution (i.e., new letters, word spellings), as the…
Optical Character RecognitionOptical Character Recognition (OCR)Parallel Spell-Checking Algorithm Based on Yahoo! N-Grams Dataset
Spell-checking is the process of detecting and sometimes providing suggestions for incorrectly spelled words in a text. Basically, the larger the dictionary of a spell-checker is, the higher is the error detection rate; …
ArticlesOCR Error Correction Using Character Correction and Feature-Based Word Classification
This paper explores the use of a learned classifier for post-OCR text correction. Experiments with the Arabic language show that this approach, which integrates a weighted confusion matrix and a shallow language model, i…
General ClassificationLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)+1A Cost Efficient Approach to Correct OCR Errors in Large Document Collections
Word error rate of an ocr is often higher than its character error rate. This is especially true when ocrs are designed by recognizing characters. High word accuracies are critical to tasks like the creation of content i…
ClusteringLanguage ModellingOptical Character Recognition (OCR)text-to-speech+1My Approach = Your Apparatus? Entropy-Based Topic Modeling on Multiple Domain-Specific Text Collections
Comparative text mining extends from genre analysis and political bias detection to the revelation of cultural and geographic differences, through to the search for prior art across patents and scientific papers. These a…
ArticlesBias DetectionClusteringDocument Classification+1