paper-with-me

홈 › Papers

Lexically Aware Semi-Supervised Learning for OCR Post-Correction

2021-11-04 · Shruti Rijhwani, Daisy Rosenblum, Antonios Anastasopoulos, Graham Neubig

Much of the existing linguistic data in many languages of the world is locked away in non-digitized books and documents. Optical character recognition (OCR) can be used to produce digitized text, and previous work has demonstrated the utility of neural post-correction methods that improve the results of general-purpose OCR systems on recognition of less-well-resourced languages. However, these methods rely on manually curated post-correction data, which are relatively scarce compared to the non-annotated raw images that need to be digitized. In this paper, we present a semi-supervised learning method that makes it possible to utilize these raw images to improve performance, specifically through the use of self-training, a technique where a model is iteratively trained on its own outputs. In addition, to enforce consistency in the recognized vocabulary, we introduce a lexically-aware decoding method that augments the neural post-correction model with a count-based language model constructed from the recognized texts, implemented using weighted finite-state automata (WFSA) for efficient and effective decoding. Results on four endangered languages demonstrate the utility of the proposed method, with relative error reductions of 15-29%, where we find the combination of self-training and lexically-aware decoding essential for achieving consistent improvements. Data and code are available at https://shrutirij.github.io/ocr-el/.

📄 PDF Abstract BibTeX arXiv:2111.02622

Code (1)

shrutirij/ocr-post-correction 공식 구현

Tasks

Language ModellingOptical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

Twice Class Bias Correction for Imbalanced Semi-Supervised Learning

2023-12-27 · Lan Li, Bowen Tao, Lu Han, De-Chuan Zhan 외

Differing from traditional semi-supervised learning, class-imbalanced semi-supervised learning presents two distinct challenges: (1) The imbalanced distribution of training samples leads to model bias towards certain cla…

Pseudo-label Correction and Learning For Semi-Supervised Object Detection

2023-03-06 · Yulin He, Wei Chen, Ke Liang, Yusong Tan 외

Pseudo-Labeling has emerged as a simple yet effective technique for semi-supervised object detection (SSOD). However, the inevitable noise problem in pseudo-labels significantly degrades the performance of SSOD methods. …

object-detectionObject DetectionPseudo LabelSemi-Supervised Object Detection

SemiGPC: Distribution-Aware Label Refinement for Imbalanced Semi-Supervised Learning Using Gaussian Processes

2023-11-03 · Abdelhak Lemkhenter, Manchen Wang, Luca Zancato, Gurumurthy Swaminathan 외

In this paper we introduce SemiGPC, a distribution-aware label refinement strategy based on Gaussian Processes where the predictions of the model are derived from the labels posterior distribution. Differently from other…

Gaussian Processes

Large Language Models Enable Few-Shot Clustering

2023-07-02 · Vijay Viswanathan, Kiril Gashteovski, Carolin Lawrence, Tongshuang Wu 외

Unlike traditional unsupervised clustering, semi-supervised clustering allows users to provide meaningful structure to the data, which helps the clustering algorithm to match the user's intent. Existing approaches to sem…

ClusteringLanguage ModelingLanguage ModellingLarge Language Model+1

Disfluency Correction using Unsupervised and Semi-supervised Learning

2021-04-01 · EACL 2021 2 · Nikhil Saini, Drumil Trivedi, Shreya Khare, Tejas Dhamecha 외

Spoken language is different from the written language in its style and structure. Disfluencies that appear in transcriptions from speech recognition systems generally hamper the performance of downstream NLP tasks. Thus…

Decoderspeech-recognitionSpeech RecognitionStyle Transfer