paper-with-me

Papers

ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images

2023-06-05 · Wenwen Yu, Chengquan Zhang, Haoyu Cao, Wei Hua, Bohan Li, Huang Chen, MingYu Liu, Mingrui Chen, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lyu, Kun Yao, Yuechen Yu, Yuliang Liu, Wanxiang Che, Errui Ding, Cheng-Lin Liu, Jiebo Luo, Shuicheng Yan, Min Zhang, Dimosthenis Karatzas, Xing Sun, Jingdong Wang, Xiang Bai

Structured text extraction is one of the most valuable and challenging application directions in the field of Document AI. However, the scenarios of past benchmarks are limited, and the corresponding evaluation protocols usually focus on the submodules of the structured text extraction scheme. In order to eliminate these problems, we organized the ICDAR 2023 competition on Structured text extraction from Visually-Rich Document images (SVRD). We set up two tracks for SVRD including Track 1: HUST-CELL and Track 2: Baidu-FEST, where HUST-CELL aims to evaluate the end-to-end performance of Complex Entity Linking and Labeling, and Baidu-FEST focuses on evaluating the performance and generalization of Zero-shot / Few-shot Structured Text extraction from an end-to-end perspective. Compared to the current document benchmarks, our two tracks of competition benchmark enriches the scenarios greatly and contains more than 50 types of visually-rich document images (mainly from the actual enterprise applications). The competition opened on 30th December, 2022 and closed on 24th March, 2023. There are 35 participants and 91 valid submissions received for Track 1, and 15 participants and 26 valid submissions received for Track 2. In this report we will presents the motivation, competition datasets, task definition, evaluation protocol, and submission summaries. According to the performance of the submissions, we believe there is still a large gap on the expected information extraction performance for complex and zero-shot scenarios. It is hoped that this competition will attract many researchers in the field of CV and NLP, and bring some new thoughts to the field of Document AI.

📄 PDF Abstract BibTeX arXiv:2306.03287

Code (0)

등록된 구현이 없습니다.

Tasks

Document AIEntity Linkingvalid

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

ICDAR2019 Competition on Scanned Receipt OCR and Information Extraction

2021-03-18 · Zheng Huang, Kai Chen, Jianhua He, Xiang Bai 외

Scanned receipts OCR and key information extraction (SROIE) represent the processeses of recognizing text from scanned receipts and extracting key texts from them and save the extracted tests to structured documents. SRO…

Key Information ExtractionOptical Character Recognition (OCR)Task 2

ICDAR 2019 Historical Document Reading Challenge on Large Structured Chinese Family Records

2019-03-08 · Rajkumar Saini, Derek Dobson, Jon Morrey, Marcus Liwicki 외

We propose a Historical Document Reading Challenge on Large Chinese Structured Family Records, in short ICDAR2019 HDRC CHINESE. The objective of the proposed competition is to recognize and analyze the layout, and finall…

System Description of CITlab's Recognition & Retrieval Engine for ICDAR2017 Competition on Information Extraction in Historical Handwritten Records

2018-04-26 · Strauß Tobias, Weidemann Max, Michael Johannes, Leifert Gundram 외

We present a recognition and retrieval system for the ICDAR2017 Competition on Information Extraction in Historical Handwritten Records which successfully infers person names and other data from marriage records. The sys…

Retrieval

ICDAR 2021 Competition on Scientific Literature Parsing

2021-06-08 · Antonio Jimeno Yepes, Xu Zhong, Douglas Burdick

Scientific literature contain important information related to cutting-edge innovations in diverse domains. Advances in natural language processing have been driving the fast development in automated information extracti…

document understandingobject-detectionObject DetectionTable Recognition

Baseline Detection in Historical Documents using Convolutional U-Nets

2018-10-22 · Michael Fink, Thomas Layer, Georg Mackenbrock, Michael Sprinzl

Baseline detection is still a challenging task for heterogeneous collections of historical documents. We present a novel approach to baseline extraction in such settings, turning out the winning entry to the ICDAR 2017 C…