paper-with-me

홈 › Papers

DocTrack: A Visually-Rich Document Dataset Really Aligned with Human Eye Movement for Machine Reading

2023-10-23 · Hao Wang, Qingxuan Wang, Yue Li, Changqing Wang, Chenhui Chu, Rui Wang

The use of visually-rich documents (VRDs) in various fields has created a demand for Document AI models that can read and comprehend documents like humans, which requires the overcoming of technical, linguistic, and cognitive barriers. Unfortunately, the lack of appropriate datasets has significantly hindered advancements in the field. To address this issue, we introduce \textsc{DocTrack}, a VRD dataset really aligned with human eye-movement information using eye-tracking technology. This dataset can be used to investigate the challenges mentioned above. Additionally, we explore the impact of human reading order on document understanding tasks and examine what would happen if a machine reads in the same order as a human. Our results suggest that although Document AI models have made significant progress, they still have a long way to go before they can read VRDs as accurately, continuously, and flexibly as humans do. These findings have potential implications for future research and development of Document AI models. The data is available at \url{https://github.com/hint-lab/doctrack}.

📄 PDF Abstract BibTeX arXiv:2310.14802

Code (1)

hint-lab/doctrack 공식 구현

Tasks

Document AIdocument understandingReading Comprehension

Similar Papers 제목 키워드 기반

DAViD: Domain Adaptive Visually-Rich Document Understanding with Synthetic Insights

2024-10-02 · Yihao Ding, Soyeon Caren Han, Zechuan Li, Hyunsuk Chung

Visually-Rich Documents (VRDs), encompassing elements like charts, tables, and references, convey complex information across various fields. However, extracting information from these rich documents is labor-intensive, e…

document understandingDomain AdaptationRepresentation Learning

VRDU: A Benchmark for Visually-rich Document Understanding

2022-11-15 · Zilong Wang, Yichao Zhou, Wei Wei, Chen-Yu Lee 외

Understanding visually-rich business documents to extract structured data and automate business workflows has been receiving attention both in academia and industry. Although recent multi-modal language models have achie…

document understanding

Document Intelligence Metrics for Visually Rich Document Evaluation

2022-05-23 · Jonathan Degange, Swapnil Gupta, Zhuoyu Han, Krzysztof Wilkosz 외

The processing of Visually-Rich Documents (VRDs) is highly important in information extraction tasks associated with Document Intelligence. We introduce DI-Metrics, a Python library devoted to VRD model evaluation compri…

Document AI

LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding

2021-04-18 · Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang 외

Multimodal pre-training with text, layout, and image has achieved SOTA performance for visually-rich document understanding tasks recently, which demonstrates the great potential for joint learning across different modal…

Document Image Classificationdocument understandingFormKey-value Pair Extraction

3MVRD: Multimodal Multi-task Multi-teacher Visually-Rich Form Document Understanding

2024-02-28 · Yihao Ding, Lorenzo Vaiani, Caren Han, Jean Lee 외

This paper presents a groundbreaking multimodal, multi-task, multi-teacher joint-grained knowledge distillation model for visually-rich form document understanding. The model is designed to leverage insights from both fi…

document understandingFormKnowledge Distillation