Evaluating Sequence-to-Sequence Models for Handwritten Text Recognition
Encoder-decoder models have become an effective approach for sequence learning tasks like machine translation, image captioning and speech recognition, but have yet to show competitive results for handwritten text recognition. To this end, we propose an attention-based sequence-to-sequence model. It combines a convolutional neural network as a generic feature extractor with a recurrent neural network to encode both the visual information, as well as the temporal context between characters in the input image, and uses a separate recurrent neural network to decode the actual character sequence. We make experimental comparisons between various attention mechanisms and positional encodings, in order to find an appropriate alignment between the input and output sequence. The model can be trained end-to-end and the optional integration of a hybrid loss allows the encoder to retain an interpretable and usable output, if desired. We achieve competitive results on the IAM and ICFHR2016 READ data sets compared to the state-of-the-art without the use of a language model, and we significantly improve over any recent sequence-to-sequence approaches.
Code (1)
Tasks
DecoderHandwritten Text RecognitionImage CaptioningKeyword SpottingLanguage ModelingLanguage ModellingMachine Translationspeech-recognitionSpeech RecognitionTranslationSimilar Papers 제목 키워드 기반
Uncovering the Handwritten Text in the Margins: End-to-end Handwritten Text Detection and Recognition
The pressing need for digitization of historical documents has led to a strong interest in designing computerised image processing methods for automatic handwritten text recognition. However, not much attention has been …
Data AugmentationHandwritten Text RecognitionText DetectionTransfer LearningContinuous Offline Handwriting Recognition using Deep Learning Models
Handwritten text recognition is an open problem of great interest in the area of automatic document image analysis. The transcription of handwritten content present in digitized documents is significant in analyzing hist…
Deep LearningHandwriting RecognitionHandwritten Text RecognitionHTRSequence-to-Sequence Contrastive Learning for Text Recognition
We propose a framework for sequence-to-sequence contrastive learning (SeqCLR) of visual representations, which we apply to text recognition. To account for the sequence-to-sequence structure, each feature map is divided …
Contrastive LearningDecoderHandwritten Text RecognitionFull Page Handwriting Recognition via Image to Sequence Extraction
We present a Neural Network based Handwritten Text Recognition (HTR) model architecture that can be trained to recognize full pages of handwritten or printed text without image segmentation. Being based on Image to Seque…
Handwriting RecognitionHandwritten Text RecognitionHTRImage Segmentation+2Candidate Fusion: Integrating Language Modelling into a Sequence-to-Sequence Handwritten Word Recognition Architecture
Sequence-to-sequence models have recently become very popular for tackling handwritten word recognition problems. However, how to effectively integrate an external language model into such recognizer is still a challengi…
Language ModelingLanguage Modelling