paper-with-me

홈 › Papers

dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model

2025-12-02 · Yumeng Li, Guang Yang, Hao Liu, Bowen Wang, Colin Zhang arxiv

Document Layout Parsing serves as a critical gateway for Artificial Intelligence (AI) to access and interpret the world's vast stores of structured knowledge. This process,which encompasses layout detection, text recognition, and relational understanding, is particularly crucial for empowering next-generation Vision-Language Models. Current methods, however, rely on fragmented, multi-stage pipelines that suffer from error propagation and fail to leverage the synergies of joint training. In this paper, we introduce dots_ocr, a single Vision-Language Model that, for the first time, demonstrates the advantages of jointly learning three core tasks within a unified, end-to-end framework. This is made possible by a highly scalable data engine that synthesizes a vast multilingual corpus, empowering the model to deliver robust performance across a wide array of tasks, encompassing diverse languages, layouts, and domains. The efficacy of our unified paradigm is validated by state-of-the-art performance on the comprehensive OmniDocBench. Furthermore, to catalyze research in global document intelligence, we introduce XDocParse, a challenging new benchmark spanning 126 languages. On this benchmark, dots_ocr achieves state-of-the-art performance, delivering an approximately 10% relative improvement and demonstrating strong multilingual capability.

📄 PDF Abstract BibTeX arXiv:2512.02498

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multimodal OCR: Parse Anything from Documents

2026-03-13 · Handong Zheng, Yumeng Li, Kaile Zhang, Liang Xin 외 arxiv

We present Multimodal OCR (MOCR), a document parsing paradigm that jointly parses text and graphics into unified textual representations. Unlike conventional OCR systems that focus on text recognition and leave graphical…

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

2026-05-31 · Minglai Yang, Xinyan Velocity Yu, Pengyuan Li, Xinyu Guo 외 arxiv

Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Optical Character Recognition (OCR) and document parsing benchmarks are i…

IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing

2025-12-23 · Oikantik Nath, Sahithi Kukkala, Mitesh Khapra, Ravi Kiran Sarvadevabhatla arxiv

Document layout analysis is essential for downstream tasks such as information retrieval, extraction, OCR, and digitization. However, existing large-scale datasets like PubLayNet and DocBank lack fine-grained region labe…

Document Layout AnalysisInformation Retrieval

RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild

2026-06-22 · Cheng Cui, Tingquan Gao, Xueqing Wang, Changda Zhou 외 arxiv

Accurate document layout analysis remains a critical bottleneck for document parsing systems, due to the intricate coupling among heterogeneous document layout elements, geometric distortions (\eg, paper warping and bend…

Document Layout Analysis

DMRST: A Joint Framework for Document-Level Multilingual RST Discourse Segmentation and Parsing

2021-10-09 · CODI 2021 11 · Zhengyuan Liu, Ke Shi, Nancy F. Chen

Text discourse parsing weighs importantly in understanding information flow and argumentative structure in natural language, making it beneficial for downstream tasks. While previous work significantly improves the perfo…

Discourse ParsingDiscourse SegmentationEnd-to-End RST ParsingSegmentation+1