Evaluation of a Region Proposal Architecture for Multi-task Document Layout Analysis
Automatically recognizing the layout of handwritten documents is an important step towards useful extraction of information from those documents. The most common application is to feed downstream applications such as automatic text recognition and keyword spotting; however, the recognition of the layout also helps to establish relationships between elements in the document which allows to enrich the information that can be extracted. Most of the modern document layout analysis systems are designed to address only one part of the document layout problem, namely: baseline detection or region segmentation. In contrast, we evaluate the effectiveness of the Mask-RCNN architecture to address the problem of baseline detection and region segmentation in an integrated manner. We present experimental results on two handwritten text datasets and one handwritten music dataset. The analyzed architecture yields promising results, outperforming state-of-the-art techniques in all three datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Document Layout AnalysisKeyword SpottingRegion ProposalSegmentationSimilar Papers 제목 키워드 기반
Fusing Saliency Maps with Region Proposals for Unsupervised Object Localization
In this paper we address the problem of unsupervised localization of objects in single images. Compared to previous state-of-the-art method our method is fully unsupervised in the sense that there is no prior instance le…
DecoderObject LocalizationRegion ProposalUnsupervised Object LocalizationRecurrent Tubelet Proposal and Recognition Networks for Action Detection
Detecting actions in videos is a challenging task as video is an information intensive media with complex variations. Existing approaches predominantly generate action proposals for each individual frame or fixed-length …
Action DetectionRegion ProposalMulti-label Image Recognition by Recurrently Discovering Attentional Regions
This paper proposes a novel deep architecture to address multi-label image recognition, a fundamental and practical task towards general visual understanding. Current solutions for this task usually rely on an extra step…
General Classificationimage-classificationImage ClassificationMulti-Label Image Classification+2Arbitrary-Oriented Scene Text Detection via Rotation Proposals
This paper introduces a novel rotation-based framework for arbitrary-oriented text detection in natural scene images. We present the Rotation Region Proposal Networks (RRPN), which are designed to generate inclined propo…
Computational EfficiencyRegion ProposalScene Text DetectionText DetectionF2DNet: Fast Focal Detection Network for Pedestrian Detection
Two-stage detectors are state-of-the-art in object detection as well as pedestrian detection. However, the current two-stage detectors are inefficient as they do bounding box regression in multiple steps i.e. in region p…
object-detectionObject DetectionPedestrian DetectionRegion Proposal