paper-with-me

홈 › Papers

MinerU: An Open-Source Solution for Precise Document Content Extraction

2024-09-27 · Bin Wang, Chao Xu, Xiaomeng Zhao, Linke Ouyang, Fan Wu, Zhiyuan Zhao, Rui Xu, Kaiwen Liu, Yuan Qu, FuKai Shang, Bo Zhang, Liqun Wei, Zhihao Sui, Wei Li, Botian Shi, Yu Qiao, Dahua Lin, Conghui He

Document content analysis has been a crucial research area in computer vision. Despite significant advancements in methods such as OCR, layout detection, and formula recognition, existing open-source solutions struggle to consistently deliver high-quality content extraction due to the diversity in document types and content. To address these challenges, we present MinerU, an open-source solution for high-precision document content extraction. MinerU leverages the sophisticated PDF-Extract-Kit models to extract content from diverse documents effectively and employs finely-tuned preprocessing and postprocessing rules to ensure the accuracy of the final results. Experimental results demonstrate that MinerU consistently achieves high performance across various document types, significantly enhancing the quality and consistency of content extraction. The MinerU open-source project is available at https://github.com/opendatalab/MinerU.

📄 PDF Abstract BibTeX arXiv:2409.18839

Code (2)

opendatalab/PDF-Extract-Kit 공식 구현 pytorch
opendatalab/mineru 공식 구현 paddle

Tasks

DiversityOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

MinerU.Chem: A High-Precision System for Optical Chemical Structure and Reaction Recognition

2026-08-04 · Haote Yang, Jiang Wu, Jingchao Wang, Xingjian Wei 외 arxiv

In organic chemistry papers and patents, molecular structures, reaction schemes, and experimental conditions are often presented as molecular structure depictions, reaction diagrams, and complex tables or figures. Such i…

Molecular Property Prediction

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

2025-09-26 · Junbo Niu, Zheng Liu, Zhuangcheng Gu, Bin Wang 외 arxiv

We introduce MinerU2.5, a 1.2B-parameter document parsing vision-language model that achieves state-of-the-art recognition accuracy while maintaining exceptional computational efficiency. Our approach employs a coarse-to…

Computational Efficiency

MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing

2026-05-24 · Bangrui Xu, Ziyang Miao, Xuanhe Zhou, Yiming Lin 외 arxiv

VLM-based OCR models have become the de facto choice for document parsing, as they can accurately extract page-level elements (e.g., paragraphs within individual pages) together with their bounding boxes and textual cont…

MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding

2026-03-23 · Hejun Dong, Junbo Niu, Bin Wang, Weijun Zeng 외 arxiv

Optical character recognition (OCR) has evolved from line-level transcription to structured document parsing, requiring models to recover long-form sequences containing layout, tables, and formulas. Despite recent advanc…

Inverse Rendering

ParseFixer: An Agentic Framework for Document Parsing via Selective Multimodal Correction

2026-06-10 · LeKai Yu, Hao Liu, Kun Wang, Zhiran Li 외 arxiv

In this report, we present our third-place solution for the DataMFM Challenge Track 1: Document Parsing. This track requires models to recover structured Markdown documents from document page images while preserving text…