VTLayout: Fusion of Visual and Text Features for Document Layout Analysis
Documents often contain complex physical structures, which make the Document Layout Analysis (DLA) task challenging. As a pre-processing step for content extraction, DLA has the potential to capture rich information in historical or scientific documents on a large scale. Although many deep-learning-based methods from computer vision have already achieved excellent performance in detecting \emph{Figure} from documents, they are still unsatisfactory in recognizing the \emph{List}, \emph{Table}, \emph{Text} and \emph{Title} category blocks in DLA. This paper proposes a VTLayout model fusing the documents' deep visual, shallow visual, and text features to localize and identify different category blocks. The model mainly includes two stages, and the three feature extractors are built in the second stage. In the first stage, the Cascade Mask R-CNN model is applied directly to localize all category blocks of the documents. In the second stage, the deep visual, shallow visual, and text features are extracted for fusion to identify the category blocks of documents. As a result, we strengthen the classification power of different category blocks based on the existing localization technique. The experimental results show that the identification capability of the VTLayout is superior to the most advanced method of DLA based on the PubLayNet dataset, and the F1 score is as high as 0.9599.
Code (0)
등록된 구현이 없습니다.
Tasks
Document Layout AnalysisMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Visual and Textual Deep Feature Fusion for Document Image Classification
The topic of text document image classification has been explored extensively over the past few years. Most recent approaches handled this task by jointly learning the visual features of document images and their corresp…
Classificationdocument-image-classificationDocument Image Classificationimage-classification+4EAML: Ensemble Self-Attention-based Mutual Learning Network for Document Image Classification
In the recent past, complex deep neural networks have received huge interest in various document understanding tasks such as document image classification and document retrieval. As many document types have a distinct vi…
document-image-classificationDocument Image Classificationimage-classificationPARL: Position-Aware Relation Learning Network for Document Layout Analysis
Document layout analysis aims to detect and categorize structural elements (e.g., titles, tables, figures) in scanned or digital documents. Popular methods often rely on high-quality Optical Character Recognition (OCR) t…
Document Layout AnalysisContent-based similar document image retrieval using fusion of CNN features
Rapid increase of digitized document give birth to high demand of document image retrieval. While conventional document image retrieval approaches depend on complex OCR-based text recognition and text similarity detectio…
Image RetrievalOptical Character Recognition (OCR)Retrievaltext similarityTowards Robust Tampered Text Detection in Document Image: New Dataset and New Solution
Recently, tampered text detection in document image has attracted increasingly attention due to its essential role on information security. However, detecting visually consistent tampered text in photographed documen…
DecoderImage and Video Forgery DetectionImage CompressionText Detection