paper-with-me

Papers

Business Document Information Extraction: Towards Practical Benchmarks

2022-06-20 · Matyáš Skalický, Štěpán Šimsa, Michal Uřičář, Milan Šulc

Information extraction from semi-structured documents is crucial for frictionless business-to-business (B2B) communication. While machine learning problems related to Document Information Extraction (IE) have been studied for decades, many common problem definitions and benchmarks do not reflect domain-specific aspects and practical needs for automating B2B document communication. We review the landscape of Document IE problems, datasets and benchmarks. We highlight the practical aspects missing in the common definitions and define the Key Information Localization and Extraction (KILE) and Line Item Recognition (LIR) problems. There is a lack of relevant datasets and benchmarks for Document IE on semi-structured business documents as their content is typically legally protected or sensitive. We discuss potential sources of available documents including synthetic data.

📄 PDF Abstract BibTeX arXiv:2206.11229

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Retrieval Augmented Structured Generation: Business Document Information Extraction As Tool Use

2024-05-30 · Franz Louis Cesista, Rui Aguiar, Jason Kim, Paolo Acilo

Business Document Information Extraction (BDIE) is the problem of transforming a blob of unstructured information (raw text, scanned documents, etc.) into a structured format that downstream systems can parse and use. It…

document understandingKey Information ExtractionLine Items Extraction

KVP10k : A Comprehensive Dataset for Key-Value Pair Extraction in Business Documents

2024-05-01 · Oshri Naparstek, Roi Pony, Inbar Shapira, Foad Abo Dahood 외

In recent years, the challenge of extracting information from business documents has emerged as a critical task, finding applications across numerous domains. This effort has attracted substantial interest from both indu…

DiversityKey Information ExtractionKey-value Pair Extraction

DocILE Benchmark for Document Information Localization and Extraction

2023-02-11 · Štěpán Šimsa, Milan Šulc, Michal Uřičář, Yash Patel 외

This paper introduces the DocILE benchmark with the largest dataset of business documents for the tasks of Key Information Localization and Extraction and Line Item Recognition. It contains 6.7k annotated business docume…

Key Information ExtractionUnsupervised Pre-training

Jointly Learning Span Extraction and Sequence Labeling for Information Extraction from Business Documents

2022-05-26 · Nguyen Hong Son, Hieu M. Vu, Tuan-Anh D. Nguyen, Minh-Tien Nguyen

This paper introduces a new information extraction model for business documents. Different from prior studies which only base on span extraction or sequence labeling, the model takes into account advantage of both span e…

Improving Information Extraction on Business Documents with Specific Pre-Training Tasks

2023-09-11 · Thibault Douzon, Stefan Duffner, Christophe Garcia, Jérémy Espinas

Transformer-based Language Models are widely used in Natural Language Processing related tasks. Thanks to their pre-training, they have been successfully adapted to Information Extraction in business documents. However, …

Language ModelingLanguage Modelling