paper-with-me

Document Layout Analysis 벤치마크

Document Layout Analysis on PubLayNet val

15개 결과 · ⬇ CSV · JSON

Overall

0.722 0.782 0.842 0.902 0.962 2019-08 2026-09 Mask RCNN — 0.91 (2019-08-16) Faster RCNN — 0.902 (2019-08-16) DeiT-B — 0.932 (2020-12-23) VSR — 0.957 (2021-05-13) BEiT-B — 0.931 (2021-06-15) DiT-L — 0.949 (2022-03-04) LayoutLMv3-B — 0.951 (2022-04-18) UDoc — 0.939 (2022-04-22) TRDLU — 0.959 (2022-10-16) DETR — 0.957 (2023-06-23) GLAM — 0.722 (2023-08-03) VGT — 0.962 (2023-08-29) ResNext-101-32×8d — 0.935 (2023-08-29) DoPTA — 0.949 (2024-12-17) Mask RCNN — 0.91 (2019-08-16) DeiT-B — 0.932 (2020-12-23) VSR — 0.957 (2021-05-13) TRDLU — 0.959 (2022-10-16) VGT — 0.962 (2023-08-29)
RankModel OverallTextTitleListTableFigure PaperCodeYear
1 VGT 0.9620.9500.9390.9680.9810.971 Vision Grid Transformer for Document Layout Analysis alibabaresearch/advancedliteratemachinery 2023
2 TRDLU 0.9590.9580.9210.9750.9760.966 Transformer-based Approach for Document Understanding 2022
3 VSR 0.9570.9670.9310.9470.9740.964 VSR: A Unified Framework for Document Layout Analysis combining Vision, Semantics and Relations hikopensource/davar-lab-ocr 2021
3 DETR 0.9570.9470.9180.9640.9810.975 Bridging the Performance Gap between DETR and R-CNN for Graphical Object Detection in Document Images 2023
5 LayoutLMv3-B 0.9510.9450.9060.9550.9790.970 LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking huggingface/transformers · microsoft/unilm · pwc-1/Paper-9 · +1 2022
6 DoPTA 0.9490.9440.8950.9570.9770.970 DoPTA: Improving Document Layout Analysis using Patch-Text Alignment 2024
6 DiT-L 0.9490.9440.8930.9600.9780.972 DiT: Self-supervised Pre-training for Document Image Transformer huggingface/transformers · microsoft/unilm · thibaultvt/Diard · +1 2022
8 UDoc 0.9390.9390.8850.9370.9730.964 Unified Pretraining Framework for Document Understanding 2022
9 ResNext-101-32×8d 0.9350.9300.8620.9400.9760.968 Vision Grid Transformer for Document Layout Analysis alibabaresearch/advancedliteratemachinery 2023
10 DeiT-B 0.9320.9340.874 0.9210.9720.957 Training data-efficient image transformers & distillation through attention huggingface/transformers · rwightman/pytorch-image-models · PaddlePaddle/PaddleClas · +37 2020
11 BEiT-B 0.9310.9340.8660.9240.973 0.957 BEiT: BERT Pre-Training of Image Transformers huggingface/transformers · rwightman/pytorch-image-models · microsoft/unilm · +11 2021
12 Mask RCNN 0.9100.9160.8400.8860.9600.949 PubLayNet: largest dataset ever for document layout analysis ibm-aur-nlp/PubLayNet · ibm-aur-nlp/PubTabNet · hpanwar08/detectron2 · +3 2019
13 Faster RCNN 0.9020.9100.8260.8830.9540.937 PubLayNet: largest dataset ever for document layout analysis ibm-aur-nlp/PubLayNet · ibm-aur-nlp/PubTabNet · hpanwar08/detectron2 · +3 2019
14 GLAM 0.7220.8780.8000.8620.8680.206 A Graphical Approach to Document Layout Analysis ivanstepanovftw/glam 2023
15 CDeC-Net 0.978 CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document Images mdv3101/CDeCNet · samarthramesh/CDeC-Net · 2023-MindSpore-4/Code2 2020
1–15 / 15 페이지당 10 20 50 100