paper
-with-
me
Papers
Browse State-of-the-Art
Datasets
Methods
AI Agents
Trends
Digest
🌙
Document Layout Analysis
벤치마크
Document Layout Analysis on PubLayNet val
15개 결과 ·
⬇ CSV
·
JSON
Overall
0.722
0.782
0.842
0.902
0.962
2019-08
2026-09
Mask RCNN — 0.91 (2019-08-16)
Faster RCNN — 0.902 (2019-08-16)
DeiT-B — 0.932 (2020-12-23)
VSR — 0.957 (2021-05-13)
BEiT-B — 0.931 (2021-06-15)
DiT-L — 0.949 (2022-03-04)
LayoutLMv3-B — 0.951 (2022-04-18)
UDoc — 0.939 (2022-04-22)
TRDLU — 0.959 (2022-10-16)
DETR — 0.957 (2023-06-23)
GLAM — 0.722 (2023-08-03)
VGT — 0.962 (2023-08-29)
ResNext-101-32×8d — 0.935 (2023-08-29)
DoPTA — 0.949 (2024-12-17)
Mask RCNN — 0.91 (2019-08-16)
DeiT-B — 0.932 (2020-12-23)
VSR — 0.957 (2021-05-13)
TRDLU — 0.959 (2022-10-16)
VGT — 0.962 (2023-08-29)
2019-08-16 — Mask RCNN: Overall 0.91
2020-12-23 — DeiT-B: Overall 0.932
2021-05-13 — VSR: Overall 0.957
2022-10-16 — TRDLU: Overall 0.959
2023-08-29 — VGT: Overall 0.962
Rank
Model
Overall
Text
Title
List
Table
Figure
Paper
Code
Year
1
VGT
0.962
0.950
0.939
0.968
0.981
0.971
Vision Grid Transformer for Document Layout Analysis
alibabaresearch/advancedliteratemachinery
2023
2
TRDLU
0.959
0.958
0.921
0.975
0.976
0.966
Transformer-based Approach for Document Understanding
2022
3
VSR
0.957
0.967
0.931
0.947
0.974
0.964
VSR: A Unified Framework for Document Layout Analysis combining Vision, Semantics and Relations
hikopensource/davar-lab-ocr
2021
3
DETR
0.957
0.947
0.918
0.964
0.981
0.975
Bridging the Performance Gap between DETR and R-CNN for Graphical Object Detection in Document Images
2023
5
LayoutLMv3-B
0.951
0.945
0.906
0.955
0.979
0.970
LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking
huggingface/transformers
·
microsoft/unilm
·
pwc-1/Paper-9
·
+1
2022
6
DoPTA
0.949
0.944
0.895
0.957
0.977
0.970
DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
2024
6
DiT-L
0.949
0.944
0.893
0.960
0.978
0.972
DiT: Self-supervised Pre-training for Document Image Transformer
huggingface/transformers
·
microsoft/unilm
·
thibaultvt/Diard
·
+1
2022
8
UDoc
0.939
0.939
0.885
0.937
0.973
0.964
Unified Pretraining Framework for Document Understanding
2022
9
ResNext-101-32×8d
0.935
0.930
0.862
0.940
0.976
0.968
Vision Grid Transformer for Document Layout Analysis
alibabaresearch/advancedliteratemachinery
2023
10
DeiT-B
0.932
0.934
0.874
0.921
0.972
0.957
Training data-efficient image transformers & distillation through attention
huggingface/transformers
·
rwightman/pytorch-image-models
·
PaddlePaddle/PaddleClas
·
+37
2020
11
BEiT-B
0.931
0.934
0.866
0.924
0.973
0.957
BEiT: BERT Pre-Training of Image Transformers
huggingface/transformers
·
rwightman/pytorch-image-models
·
microsoft/unilm
·
+11
2021
12
Mask RCNN
0.910
0.916
0.840
0.886
0.960
0.949
PubLayNet: largest dataset ever for document layout analysis
ibm-aur-nlp/PubLayNet
·
ibm-aur-nlp/PubTabNet
·
hpanwar08/detectron2
·
+3
2019
13
Faster RCNN
0.902
0.910
0.826
0.883
0.954
0.937
PubLayNet: largest dataset ever for document layout analysis
ibm-aur-nlp/PubLayNet
·
ibm-aur-nlp/PubTabNet
·
hpanwar08/detectron2
·
+3
2019
14
GLAM
0.722
0.878
0.800
0.862
0.868
0.206
A Graphical Approach to Document Layout Analysis
ivanstepanovftw/glam
2023
15
CDeC-Net
–
–
–
–
0.978
–
CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document Images
mdv3101/CDeCNet
·
samarthramesh/CDeC-Net
·
2023-MindSpore-4/Code2
2020
1–15 / 15
페이지당
10
20
50
100