paper-with-me

Papers

Document Layout Annotation: Database and Benchmark in the Domain of Public Affairs

2023-06-12 · Alejandro Peña, Aythami Morales, Julian Fierrez, Javier Ortega-Garcia, Marcos Grande, Iñigo Puente, Jorge Cordova, Gonzalo Cordova

Every day, thousands of digital documents are generated with useful information for companies, public organizations, and citizens. Given the impossibility of processing them manually, the automatic processing of these documents is becoming increasingly necessary in certain sectors. However, this task remains challenging, since in most cases a text-only based parsing is not enough to fully understand the information presented through different components of varying significance. In this regard, Document Layout Analysis (DLA) has been an interesting research field for many years, which aims to detect and classify the basic components of a document. In this work, we used a procedure to semi-automatically annotate digital documents with different layout labels, including 4 basic layout blocks and 4 text categories. We apply this procedure to collect a novel database for DLA in the public affairs domain, using a set of 24 data sources from the Spanish Administration. The database comprises 37.9K documents with more than 441K document pages, and more than 8M labels associated to 8 layout block units. The results of our experiments validate the proposed text labeling procedure with accuracy up to 99%.

📄 PDF Abstract BibTeX arXiv:2306.10046

Code (1)

bidalab/paldb 공식 구현

Tasks

Document Layout Analysis

Methods 이 논문이 사용한 방법론

DLA 설명 없음

Similar Papers 제목 키워드 기반

WordScape: a Pipeline to extract multilingual, visually rich Documents with Layout Annotations from Web Crawl Data

2023-12-15 · NeurIPS 2023 11 · Maurice Weber, Carlo Siebenschuh, Rory Butler, Anton Alexandrov 외

We introduce WordScape, a novel pipeline for the creation of cross-disciplinary, multilingual corpora comprising millions of pages with annotations for document layout detection. Relating visual and textual items on docu…

document understandingQuestion AnsweringVisual Question Answering

BaDLAD: A Large Multi-Domain Bengali Document Layout Analysis Dataset

2023-03-09 · Md. Istiak Hossain Shihab, Md. Rakibul Hasan, Mahfuzur Rahman Emon, Syed Mobassir Hossen 외

While strides have been made in deep learning based Bengali Optical Character Recognition (OCR) in the past decade, the absence of large Document Layout Analysis (DLA) datasets has hindered the application of OCR in docu…

BenchmarkingDeep LearningDocument Layout AnalysisOptical Character Recognition+1

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

2026-05-31 · Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo 외 arxiv

Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are i…

DocILE Benchmark for Document Information Localization and Extraction

2023-02-11 · Štěpán Šimsa, Milan Šulc, Michal Uřičář, Yash Patel 외

This paper introduces the DocILE benchmark with the largest dataset of business documents for the tasks of Key Information Localization and Extraction and Line Item Recognition. It contains 6.7k annotated business docume…

Key Information ExtractionUnsupervised Pre-training

Cross-Domain Document Object Detection: Benchmark Suite and Method

2020-03-30 · CVPR 2020 6 · Kai Li, Curtis Wigington, Chris Tensmeyer, Handong Zhao 외

Decomposing images of document pages into high-level semantic regions (e.g., figures, tables, paragraphs), document object detection (DOD) is fundamental for downstream tasks like intelligent document editing and underst…

object-detectionObject Detection