paper-with-me

홈 › Papers

How Far Is Document Parsing from Solved? PureDocBench: A Source-TraceableBenchmark across Clean, Degraded, and Real-World Settings

2026-05-08 · Zhiheng Li, Zongyang Ma, Jiaxian Chen, Jianing Zhang, Zhaolong Su, Yutong Zhang, Zhiyin Yu, Ruiqi Liu, Xiaolei Lv, Bo Li, Jun Gao, Ziqi Zhang, Chunfeng Yuan, Bing Li, Weiming Hu arxiv

The past year has seen over 20 open-source document parsing models, yet thefield still benchmarks almost exclusively on OmniDocBench, a 1,355-pagemanually annotated dataset whose top scores have saturated above 90%. Athree-stage audit pipeline we run on OmniDocBench screens its 21,353evaluator-scored blocks and confirms 2,580 errors (12.08%); combined with overa year of public availability, both annotation quality and contamination riskcall its rankings into question. To address these issues, we presentPureDocBench, a programmatically generated, source-traceable benchmark thatrenders document images from HTML/CSS and produces verifiable annotations fromthe same source, covering 10 domains, 66 subcategories, and 1,475 pages, eachin three versions: clean, digitally degraded, and real-degraded (4,425 imagestotal). Evaluating 40 models spanning pipeline specialists, end-to-endspecialists, and general-purpose VLMs, we find: (i) document parsing is farfrom solved: the best model scores only ~74 out of 100, with a 44.6-point gapbetween the strongest and weakest models; (ii) specialist parsers with <=4Bparameters rival or surpass general VLMs that are 5-100x larger, yet formularecognition remains a shared bottleneck where no model exceeds 67% whenaveraging the formula metric across all three tracks; (iii) general VLMs loseonly 0.99/8.52 Overall points under digital/real degradation versus 4.90/14.21for pipeline specialists, producing ranking reversals that make clean-onlyevaluation misleading for deployment. All data, code, and artifacts arepublicly released.

📄 PDF Abstract BibTeX arXiv:2605.07492

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

2026-09-17 · Hao Yu, Kang Liu, Linnan Zhao, Jiabo Zhan 외 hf

Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clea…

OvisOCR2 Technical Report

2026-07-15 · Shiyin Lu, Yinglun Li, Yu Xia, Yuhui Chen 외 hf

We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdown representation in natural reading order, covering text, formulas, t…

Reinforcement Learning

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

2026-08-13 · Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang 외 arxiv

Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, exis…

Representation Learning

MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios

2026-03-30 · Zhang Li, Zhibo Lin, Qiang Liu, Ziyang Zhang 외 arxiv

We introduce Multilingual Document Parsing Benchmark, the first benchmark for multilingual digital and photographed document parsing. Document parsing has made remarkable strides, yet almost exclusively on clean, digital…

Slide, Constrain, Parse, Repeat: Synchronous SlidingWindows for Document AMR Parsing

2023-05-26 · Sadhana Kumaravel, Tahira Naseem, Ramon Fernandez Astudillo, Radu Florian 외

The sliding window approach provides an elegant way to handle contexts of sizes larger than the Transformer's input window, for tasks like language modeling. Here we extend this approach to the sequence-to-sequence task …

Abstract Meaning RepresentationAMR ParsingLanguage ModelingLanguage Modelling+1