paper-with-me

Papers

LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis

2025-07-31 · Inbum Heo, Taewook Hwang, Jeesu Jung, Sangkeun Jung arxiv

Recent advancements in Document Layout Analysis through Large Language Models and Multimodal Models have significantly improved layout detection. However, despite these improvements, challenges remain in addressing critical structural errors, such as region merging, splitting, and missing content. Conventional evaluation metrics like IoU and mAP, which focus primarily on spatial overlap, are insufficient for detecting these errors. To address this limitation, we propose Layout Error Detection (LED), a novel benchmark designed to evaluate the structural robustness of document layout predictions. LED defines eight standardized error types, and formulates three complementary tasks: error existence detection, error type classification, and element-wise error type classification. Furthermore, we construct LED-Dataset, a synthetic dataset generated by injecting realistic structural errors based on empirical distributions from DLA models. Experimental results across a range of LMMs reveal that LED effectively differentiates structural understanding capabilities, exposing modality biases and performance trade-offs not visible through traditional metrics.

📄 PDF Abstract BibTeX arXiv:2507.23295

Code (0)

등록된 구현이 없습니다.

Tasks

Document Layout Analysis

Similar Papers 제목 키워드 기반

LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis

2026-03-18 · Inbum Heo, Taewook Hwang, Jeesu Jung, Sangkeun Jung arxiv

Recent advances in Large Language Models (LLMs) and Large Multimodal Models (LMMs) have improved Document Layout Analysis (DLA), yet structural errors such as region merging, splitting, and omission remain persistent. Co…

Document Layout Analysis

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

2026-05-31 · Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo 외 arxiv

Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are i…

Parser-Oriented Structural Refinement for a Stable Layout Interface in Document Parsing

2026-04-03 · Fuyuan Liu, Dianyu Yu, He Ren, Nayu Liu 외 arxiv

Accurate document parsing requires both robust content recognition and a stable parser interface. In explicit Document Layout Analysis (DLA) pipelines, downstream parsers do not consume the full detector output. Instead,…

Document Layout Analysis

A Survey of OCR Evaluation Methods and Metrics and the Invisibility of Historical Documents

2026-03-26 · Fitsum Sileshi Beyene, Christopher L. Dancy arxiv

Optical character recognition (OCR) and document understanding systems increasingly rely on large vision and vision-language models, yet evaluation remains centered on modern, Western, and institutional documents. This e…

PARL: Position-Aware Relation Learning Network for Document Layout Analysis

2026-01-12 · Fuyuan Liu, Dianyu Yu, He Ren, Nayu Liu 외 arxiv

Document layout analysis aims to detect and categorize structural elements (e.g., titles, tables, figures) in scanned or digital documents. Popular methods often rely on high-quality Optical Character Recognition (OCR) t…

Document Layout Analysis