paper-with-me

Papers

Document Intelligence Metrics for Visually Rich Document Evaluation

2022-05-23 · Jonathan Degange, Swapnil Gupta, Zhuoyu Han, Krzysztof Wilkosz, Adam Karwan

The processing of Visually-Rich Documents (VRDs) is highly important in information extraction tasks associated with Document Intelligence. We introduce DI-Metrics, a Python library devoted to VRD model evaluation comprising text-based, geometric-based and hierarchical metrics for information extraction tasks. We apply DI-Metrics to evaluate information extraction performance using publicly available CORD dataset, comparing performance of three SOTA models and one industry model. The open-source library is available on GitHub.

📄 PDF Abstract BibTeX arXiv:2205.11215

Code (1)

metricsdi/dimetrics 공식 구현

Tasks

Document AI

Similar Papers 제목 키워드 기반

DISCO: Document Intelligence Suite for COmparative Evaluation

2026-03-04 · Kenza Benkirane, Dan Goldwater, Martin Asenov, Aneiss Ghodsi arxiv

Document intelligence requires accurate text extraction and reliable reasoning over document content. We introduce \textbf{DISCO}, a \emph{Document Intelligence Suite for COmparative Evaluation}, that evaluates optical c…

Question Answering

Table Detection for Visually Rich Document Images

2023-05-30 · Bin Xiao, Murat Simsek, Burak Kantarci, Ala Abu Alkheir

Table Detection (TD) is a fundamental task to enable visually rich document understanding, which requires the model to extract information without information loss. However, popular Intersection over Union (IoU) based ev…

document understandingobject-detectionObject DetectionPrediction+1

Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval

2026-02-23 · Yibo Yan, Jiahao Huo, Guanbo Feng, Mingdong Ou 외 arxiv

With the rapid proliferation of multimodal information, Visual Document Retrieval (VDR) has emerged as a critical frontier in bridging the gap between unstructured visually rich data and precise information acquisition. …

Image Retrieval

GraphRevisedIE: Multimodal Information Extraction with Graph-Revised Network

2024-10-02 · Panfeng Cao, Jian Wu

Key information extraction (KIE) from visually rich documents (VRD) has been a challenging task in document intelligence because of not only the complicated and diverse layouts of VRD that make the model hard to generali…

Key Information Extraction

VRDU: A Benchmark for Visually-rich Document Understanding

2022-11-15 · Zilong Wang, Yichao Zhou, Wei Wei, Chen-Yu Lee 외

Understanding visually-rich business documents to extract structured data and automate business workflows has been receiving attention both in academia and industry. Although recent multi-modal language models have achie…

document understanding