paper-with-me

홈 › Papers

READoc: A Unified Benchmark for Realistic Document Structured Extraction

2024-09-08 · Zichao Li, Aizier Abulaiti, Yaojie Lu, Xuanang Chen, Jia Zheng, Hongyu Lin, Xianpei Han, Le Sun

Document Structured Extraction (DSE) aims to extract structured content from raw documents. Despite the emergence of numerous DSE systems, their unified evaluation remains inadequate, significantly hindering the field's advancement. This problem is largely attributed to existing benchmark paradigms, which exhibit fragmented and localized characteristics. To address these limitations and offer a thorough evaluation of DSE systems, we introduce a novel benchmark named READoc, which defines DSE as a realistic task of converting unstructured PDFs into semantically rich Markdown. The READoc dataset is derived from 2,233 diverse and real-world documents from arXiv and GitHub. In addition, we develop a DSE Evaluation S$^3$uite comprising Standardization, Segmentation and Scoring modules, to conduct a unified evaluation of state-of-the-art DSE approaches. By evaluating a range of pipeline tools, expert visual models, and general VLMs, we identify the gap between current work and the unified, realistic DSE objective for the first time. We aspire that READoc will catalyze future research in DSE, fostering more comprehensive and practical solutions.

📄 PDF Abstract BibTeX arXiv:2409.05137

Code (1)

icip-cas/READoc 공식 구현 pytorch

Similar Papers 제목 키워드 기반

MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing

2026-05-21 · Bangbang Zhou, Hangdi Xing, Yifan Chen, Jianjun Xu 외 arxiv

Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information systems. Although many benchmarks have been proposed for document parsing, …

KoViDoRe: Korean Visual Document Retrieval

2026-08-21 · Yongbin Choi, Yongwoo Song, Mujeen Sung arxiv

Recent advances in multimodal retrieval have improved the ability to retrieve information from visually rich documents such as PDFs and reports. However, existing benchmarks remain largely centered on English and provide…

Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents

2026-07-01 · Ádám Kovács, Bowei He, Xue Liu, István Boros 외 arxiv

Hallucination detection for retrieval-augmented generation (RAG) is usually evaluated on natural-language document evidence. However, grounded generation systems increasingly rely on structured inputs: source code, devel…

Large Language Models for Analyzing Enterprise Architecture Debt in Unstructured Documentation

2026-03-29 · Christin Pagels, Simon Hacks, Rob Henk Bemthuis arxiv

Enterprise Architecture Debt (EA Debt) arises from suboptimal design decisions and misaligned components that can degrade an organization's IT landscape over time. Early indicators, Enterprise Architecture Smells (EA Sme…

Benchmarking Large Language Models on Reference Extraction and Parsing in the Social Sciences and Humanities

2026-03-13 · Yurui Zhu, Giovanni Colavizza, Matteo Romanello arxiv

Bibliographic reference extraction and parsing are foundational for citation indexing, linking, and downstream scholarly knowledge-graph construction. However, most established evaluations focus on clean, English, end-of…