paper-with-me

Papers

DocILE 2023 Teaser: Document Information Localization and Extraction

2023-01-29 · Štěpán Šimsa, Milan Šulc, Matyáš Skalický, Yash Patel, Ahmed Hamdi

The lack of data for information extraction (IE) from semi-structured business documents is a real problem for the IE community. Publications relying on large-scale datasets use only proprietary, unpublished data due to the sensitive nature of such documents. Publicly available datasets are mostly small and domain-specific. The absence of a large-scale public dataset or benchmark hinders the reproducibility and cross-evaluation of published methods. The DocILE 2023 competition, hosted as a lab at the CLEF 2023 conference and as an ICDAR 2023 competition, will run the first major benchmark for the tasks of Key Information Localization and Extraction (KILE) and Line Item Recognition (LIR) from business documents. With thousands of annotated real documents from open sources, a hundred thousand of generated synthetic documents, and nearly a million unlabeled documents, the DocILE lab comes with the largest publicly available dataset for KILE and LIR. We are looking forward to contributions from the Computer Vision, Natural Language Processing, Information Retrieval, and other communities. The data, baselines, code and up-to-date information about the lab and competition are available at https://docile.rossum.ai/.

📄 PDF Abstract BibTeX arXiv:2301.12394

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

DocILE Benchmark for Document Information Localization and Extraction

2023-02-11 · Štěpán Šimsa, Milan Šulc, Michal Uřičář, Yash Patel 외

This paper introduces the DocILE benchmark with the largest dataset of business documents for the tasks of Key Information Localization and Extraction and Line Item Recognition. It contains 6.7k annotated business docume…

Key Information ExtractionUnsupervised Pre-training

TeaserGen: Generating Teasers for Long Documentaries

2024-10-08 · Weihan Xu, Paul Pu Liang, Haven Kim, Julian McAuley 외

Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling…

Language ModellingLarge Language Model

Retrieval Augmented Structured Generation: Business Document Information Extraction As Tool Use

2024-05-30 · Franz Louis Cesista, Rui Aguiar, Jason Kim, Paolo Acilo

Business Document Information Extraction (BDIE) is the problem of transforming a blob of unstructured information (raw text, scanned documents, etc.) into a structured format that downstream systems can parse and use. It…

document understandingKey Information ExtractionLine Items Extraction

QUEST: Quality-aware Semi-supervised Table Extraction for Business Documents

2025-06-17 · Eliott Thomas, Mickael Coustaty, Aurelie Joseph, Gaspar Deloin 외

Automating table extraction (TE) from business documents is critical for industrial workflows but remains challenging due to sparse annotations and error-prone multi-stage pipelines. While semi-supervised learning (SSL) …

Pseudo LabelTable Extraction

PrIeD-KIE: Towards Privacy Preserved Document Key Information Extraction

2023-10-05 · Saifullah Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed

In this paper, we introduce strategies for developing private Key Information Extraction (KIE) systems by leveraging large pretrained document foundation models in conjunction with differential privacy (DP), federated le…

Document AIFederated LearningKey Information Extraction