paper-with-me

홈 › Papers

WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

2026-09-17 · Hao Yu, Kang Liu, Linnan Zhao, Jiabo Zhan, Chong Sun, Chen Li, Jing Lyu hf

Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser's remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis. Stage II uses a held-out probe to measure the Stage I parser's residual errors within fixed visual-structural clusters. These diagnostics guide targeted data construction and reallocation of the target-token budget. WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Compared with Stage I, Stage II improves Overall scores for the 2B and 4B models on both benchmarks, with larger gains on the degraded PureDocBench tracks, including a 4.03-point gain for the 4B model on the Real Degraded track.

📄 PDF Abstract BibTeX arXiv:2609.20423

Code (5)

InsomaniacElf/sg-tamil-tts-resources- ★ 1
Tavish9/awesome-daily-AI-arxiv ★ 120
Tencent/WeVisDoc ★ 42
🤗 tencent/WeVisDoc-2B ★ 9
🤗 tencent/WeVisDoc-4B ★ 26

Similar Papers 제목 키워드 기반

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

2026-07-06 · Gengluo Li, Xingyu Wan, Shangpin Peng, Weinong Wang 외 arxiv

We present HunyuanOCR-1.5, a lightweight end-to-end OCR-specialized vision-language model. HunyuanOCR unifies document parsing, text spotting, information extraction, text-image translation, and multi-image document unde…

Information ExtractionText Spotting

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

2026-05-31 · Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo 외 arxiv

Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are i…

OCRTurk: A Comprehensive OCR Benchmark for Turkish

2026-02-03 · Deniz Yılmaz, Evren Ayberk Munis, Çağrı Toraman, Süha Kağan Köse 외 arxiv

Document parsing is now widely used in applications, such as large-scale document digitization, retrieval-augmented generation, and domain-specific pipelines in healthcare and education. Benchmarking these models is cruc…

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

2026-08-13 · Peng Cai, Zhaofan Zou, Shifa Liu, Yikun Wang 외 arxiv

Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, exis…

Representation Learning

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

2024-12-10 · CVPR 2025 1 · Linke Ouyang, Yuan Qu, Hongbin Zhou, Jiawei Zhu 외

Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems. Despite recent progress, current document p…

AttributeBenchmarkingDiversityRAG+1