paper-with-me

홈 › Papers

DLAFormer: An End-to-End Transformer For Document Layout Analysis

2024-05-20 · Jiawei Wang, Kai Hu, Qiang Huo

Document layout analysis (DLA) is crucial for understanding the physical layout and logical structure of documents, serving information retrieval, document summarization, knowledge extraction, etc. However, previous studies have typically used separate models to address individual sub-tasks within DLA, including table/figure detection, text region detection, logical role classification, and reading order prediction. In this work, we propose an end-to-end transformer-based approach for document layout analysis, called DLAFormer, which integrates all these sub-tasks into a single model. To achieve this, we treat various DLA sub-tasks (such as text region detection, logical role classification, and reading order prediction) as relation prediction problems and consolidate these relation prediction labels into a unified label space, allowing a unified relation prediction module to handle multiple tasks concurrently. Additionally, we introduce a novel set of type-wise queries to enhance the physical meaning of content queries in DETR. Moreover, we adopt a coarse-to-fine strategy to accurately identify graphical page objects. Experimental results demonstrate that our proposed DLAFormer outperforms previous approaches that employ multi-branch or multi-stage architectures for multiple tasks on two document layout analysis benchmarks, DocLayNet and Comp-HRDoc.

📄 PDF Abstract BibTeX arXiv:2405.11757

Code (0)

등록된 구현이 없습니다.

Tasks

Document Layout AnalysisDocument SummarizationInformation RetrievalPredictionRelationRelation Prediction

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
DLA 설명 없음

Similar Papers 제목 키워드 기반

Document AI: A Comparative Study of Transformer-Based, Graph-Based Models, and Convolutional Neural Networks For Document Layout Analysis

2023-08-29 · Sotirios Kastanas, Shaomu Tan, Yi He

Document AI aims to automatically analyze documents by leveraging natural language processing and computer vision techniques. One of the major tasks of Document AI is document layout analysis, which structures document p…

Document AIDocument Layout AnalysisMachine TranslationTransfer Learning

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

2026-06-22 · Cheng Cui, Tingquan Gao, Xueqing Wang, Changda Zhou 외 arxiv

Accurate document layout analysis remains a critical bottleneck for document parsing systems, due to the intricate coupling among heterogeneous document layout elements, geometric distortions (\eg, paper warping and bend…

Document Layout Analysis

M6Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis

2023-01-01 · CVPR 2023 1 · Hiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang 외

Document layout analysis is a crucial prerequisite for document understanding, including document retrieval and conversion. Most public datasets currently contain only PDF documents and lack realistic documents. Mode…

ArticlesDocument Layout Analysisdocument understandingInstance Segmentation+2

M$^{6}$Doc: A Large-Scale Multi-Format, Multi-Type, Multi-Layout, Multi-Language, Multi-Annotation Category Dataset for Modern Document Layout Analysis

2023-05-15 · Hiuyi Cheng, Peirong Zhang, Sihang Wu, Jiaxin Zhang 외

Document layout analysis is a crucial prerequisite for document understanding, including document retrieval and conversion. Most public datasets currently contain only PDF documents and lack realistic documents. Models t…

ArticlesDocument Layout Analysisdocument understandingInstance Segmentation+2

Vision Grid Transformer for Document Layout Analysis

2023-08-29 · ICCV 2023 1 · Cheng Da, Chuwei Luo, Qi Zheng, Cong Yao

Document pre-trained models and grid-based models have proven to be very effective on various tasks in Document AI. However, for the document layout analysis (DLA) task, existing document pre-trained models, even those p…

Document AIDocument Layout Analysisdocument understandingOptical Character Recognition (OCR)