paper-with-me

홈 › Papers

XDoc: Unified Pre-training for Cross-Format Document Understanding

2022-10-06 · Jingye Chen, Tengchao Lv, Lei Cui, Cha Zhang, Furu Wei

The surge of pre-training has witnessed the rapid development of document understanding recently. Pre-training and fine-tuning framework has been effectively used to tackle texts in various formats, including plain texts, document texts, and web texts. Despite achieving promising performance, existing pre-trained models usually target one specific document format at one time, making it difficult to combine knowledge from multiple document formats. To address this, we propose XDoc, a unified pre-trained model which deals with different document formats in a single model. For parameter efficiency, we share backbone parameters for different formats such as the word embedding layer and the Transformer layers. Meanwhile, we introduce adaptive layers with lightweight parameters to enhance the distinction across different formats. Experimental results have demonstrated that with only 36.7% parameters, XDoc achieves comparable or even better performance on a variety of downstream tasks compared with the individual pre-trained models, which is cost effective for real-world deployment. The code and pre-trained models will be publicly available at \url{https://aka.ms/xdoc}.

📄 PDF Abstract BibTeX arXiv:2210.02849

Code (1)

microsoft/unilm/tree/master/xdoc 공식 구현 pytorch

Tasks

document understandingSemantic entity labeling

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Multi-Head Attention 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models

2025-10-02 · Karan Dua, Hitesh Laxmichand Patel, Puneet Mittal, Ranjeet Gupta 외 arxiv

Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such data is prohibitively expensive due to p…

Key Information ExtractionSynthetic Data Generation

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

2026-08-24 · Hang Wang, Jin Zhang, Guoliang Xu, Pengyue Lu 외 arxiv

Financial document parsing requires accuracy, structural consistency, and verifiability that current benchmarks often fail to reflect. We present FinixDoc, an end-to-end agentic parsing system for real-world financial do…

Reinforcement LearningContrastive Learning

Pre-Training on Large-Scale Generated Docking Conformations with HelixDock to Unlock the Potential of Protein-ligand Structure Prediction Models

2023-10-21 · Lihang Liu, Shanzhuo Zhang, Donglong He, Xianbin Ye 외

Protein-ligand structure prediction is an essential task in drug discovery, predicting the binding interactions between small molecules (ligands) and target proteins (receptors). Recent advances have incorporated deep le…

CPUDrug DiscoveryMolecular DockingPrediction

dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model

2025-12-02 · Yumeng Li, Guang Yang, Hao Liu, Bowen Wang 외 arxiv

Document Layout Parsing serves as a critical gateway for Artificial Intelligence (AI) to access and interpret the world's vast stores of structured knowledge. This process,which encompasses layout detection, text recogni…

Knowledge-Driven Cross-Document Relation Extraction

2024-05-22 · Monika Jain, Raghava Mutharaju, Kuldeep Singh, Ramakanth Kavuluru

Relation extraction (RE) is a well-known NLP application often treated as a sentence- or document-level task. However, a handful of recent efforts explore it across documents or in the cross-document setting (CrossDocRE)…

RelationRelation ExtractionSentence