paper-with-me

Papers

VRDU: A Benchmark for Visually-rich Document Understanding

2022-11-15 · Zilong Wang, Yichao Zhou, Wei Wei, Chen-Yu Lee, Sandeep Tata

Understanding visually-rich business documents to extract structured data and automate business workflows has been receiving attention both in academia and industry. Although recent multi-modal language models have achieved impressive results, we find that existing benchmarks do not reflect the complexity of real documents seen in industry. In this work, we identify the desiderata for a more comprehensive benchmark and propose one we call Visually Rich Document Understanding (VRDU). VRDU contains two datasets that represent several challenges: rich schema including diverse data types as well as hierarchical entities, complex templates including tables and multi-column layouts, and diversity of different layouts (templates) within a single document type. We design few-shot and conventional experiment settings along with a carefully designed matching algorithm to evaluate extraction results. We report the performance of strong baselines and offer three observations: (1) generalizing to new document templates is still very challenging, (2) few-shot performance has a lot of headroom, and (3) models struggle with hierarchical fields such as line-items in an invoice. We plan to open source the benchmark and the evaluation toolkit. We hope this helps the community make progress on these challenging tasks in extracting structured data from visually rich documents.

📄 PDF Abstract BibTeX arXiv:2211.15421

Code (0)

등록된 구현이 없습니다.

Tasks

document understanding

Similar Papers 제목 키워드 기반

ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training

2024-10-14 · Zhouqiang Jiang, Bowen Wang, JunHao Chen, Yuta Nakashima

Recent approaches for visually-rich document understanding (VrDU) uses manually annotated semantic groups, where a semantic group encompasses all semantically relevant but not obviously grouped words. As OCR tools are un…

document understandingOptical Character Recognition (OCR)

Deep Learning based Visually Rich Document Content Understanding: A Survey

2024-08-02 · Yihao Ding, Jean Lee, Soyeon Caren Han

Visually Rich Documents (VRDs) are essential in academia, finance, medical fields, and marketing due to their multimodal information content. Traditional methods for extracting information from VRDs depend on expert know…

Deep Learningdocument understandingMarketingSurvey

Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding

2022-06-27 · Chuwei Luo, Guozhi Tang, Qi Zheng, Cong Yao 외

Multi-modal document pre-trained models have proven to be very effective in a variety of visually-rich document understanding (VrDU) tasks. Though existing document pre-trained models have achieved excellent performance …

Document Classificationdocument understandingLanguage ModelingLanguage Modelling+1

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

2025-01-04 · Camille Barboule, Benjamin Piwowarski, Yoan Chabot

Using Large Language Models (LLMs) for Visually-rich Document Understanding (VrDU) has significantly improved performance on tasks requiring both comprehension and generation, such as question answering, albeit introduci…

document understandingQuestion AnsweringSurvey

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends

2025-07-14 · Yihao Ding, Siwen Luo, Yue Dai, Yanbei Jiang 외

Visually-Rich Document Understanding (VRDU) has emerged as a critical field, driven by the need to automatically process documents containing complex visual, textual, and layout information. Recently, Multimodal Large La…

document understandingOptical Character RecognitionOptical Character Recognition (OCR)