Annotation-free Learning of Deep Representations for Word Spotting using Synthetic Data and Self Labeling
Word spotting is a popular tool for supporting the first exploration of historic, handwritten document collections. Today, the best performing methods rely on machine learning techniques, which require a high amount of annotated training material. As training data is usually not available in the application scenario, annotation-free methods aim at solving the retrieval task without representative training samples. In this work, we present an annotation-free method that still employs machine learning techniques and therefore outperforms other learning-free approaches. The weakly supervised training scheme relies on a lexicon, that does not need to precisely fit the dataset. In combination with a confidence based selection of pseudo-labeled training samples, we achieve state-of-the-art query-by-example performances. Furthermore, our method allows to perform query-by-string, which is usually not the case for other annotation-free methods.
Code (0)
등록된 구현이 없습니다.
Tasks
BIG-bench Machine LearningRetrievalSimilar Papers 제목 키워드 기반
Learning Deep Representations for Word Spotting Under Weak Supervision
Convolutional Neural Networks have made their mark in various fields of computer vision in recent years. They have achieved state-of-the-art performance in the field of document analysis as well. However, CNNs require a …
Word Spotting In Handwritten DocumentsBootstrapping Weakly Supervised Segmentation-free Word Spotting through HMM-based Alignment
Recent work in word spotting in handwritten documents has yielded impressive results. This progress has largely been made by supervised learning systems, which are dependent on manually annotated data, making deployment …
Weakly supervised segmentationWord Spotting In Handwritten DocumentsAraSpot: Arabic Spoken Command Spotting
Spoken keyword spotting (KWS) is the task of identifying a keyword in an audio stream and is widely used in smart devices at the edge in order to activate voice assistants and perform hands-free tasks. The task is daunti…
Data AugmentationKeyword SpottingSynthetic Data Generationtext-to-speech+1Radial Line Fourier Descriptor for Historical Handwritten Text Representation
Automatic recognition of historical handwritten manuscripts is a daunting task due to paper degradation over time. Recognition-free retrieval or word spotting is popularly used for information retrieval and digitization …
Information RetrievalRetrievalR-PHOC: Segmentation-Free Word Spotting using CNN
This paper proposes a region based convolutional neural network for segmentation-free word spotting. Our net- work takes as input an image and a set of word candidate bound- ing boxes and embeds all bounding boxes into a…
Segmentation