paper-with-me

Papers

CORU: Comprehensive Post-OCR Parsing and Receipt Understanding Dataset

2024-06-06 · Abdelrahman Abdallah, Mahmoud Abdalla, Mahmoud SalahEldin Kasem, Mohamed Mahmoud, Ibrahim Abdelhalim, Mohamed Elkasaby, Yasser Elbendary, Adam Jatowt

In the fields of Optical Character Recognition (OCR) and Natural Language Processing (NLP), integrating multilingual capabilities remains a critical challenge, especially when considering languages with complex scripts such as Arabic. This paper introduces the Comprehensive Post-OCR Parsing and Receipt Understanding Dataset (CORU), a novel dataset specifically designed to enhance OCR and information extraction from receipts in multilingual contexts involving Arabic and English. CORU consists of over 20,000 annotated receipts from diverse retail settings, including supermarkets and clothing stores, alongside 30,000 annotated images for OCR that were utilized to recognize each detected line, and 10,000 items annotated for detailed information extraction. These annotations capture essential details such as merchant names, item descriptions, total prices, receipt numbers, and dates. They are structured to support three primary computational tasks: object detection, OCR, and information extraction. We establish the baseline performance for a range of models on CORU to evaluate the effectiveness of traditional methods, like Tesseract OCR, and more advanced neural network-based approaches. These baselines are crucial for processing the complex and noisy document layouts typical of real-world receipts and for advancing the state of automated multilingual document processing. Our datasets are publicly accessible (https://github.com/Update-For-Integrated-Business-AI/CORU).

📄 PDF Abstract BibTeX arXiv:2406.04493

Code (1)

update-for-integrated-business-ai/coru 공식 구현 pytorch

Tasks

object-detectionObject DetectionOptical Character RecognitionOptical Character Recognition (OCR)

Similar Papers 제목 키워드 기반

CORD: A Consolidated Receipt Dataset for Post-OCR Parsing

2019-09-14 · NeurIPS Workshop Document_Intelligen 2019 12 · Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee 외

OCR is inevitably linked to NLP since its final output is in text. Advances in document intelligence are driving the need for a unified technology that integrates OCR with various NLP tasks, especially semantic parsing. …

Optical Character Recognition (OCR)Semantic Parsing

Post-OCR parsing: building simple and robust parser via BIO tagging

2019-09-14 · NeurIPS Workshop Document_Intelligen 2019 12 · Wonseok Hwang, Seonghyeon Kim, Minjoon Seo, Jinyeong Yim 외

Parsing textual information embedded in images is important for various down- stream tasks. However, many previously developed parsers are limited to handling the information presented in one dimensional sequence format.…

Optical Character RecognitionOptical Character Recognition (OCR)

From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding

2026-05-21 · Yandi Wang, Libin Zhan, Ziwei Huang, Tiancheng Luo 외 arxiv

Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multimodal Large Language Models (MLLMs) have shown promising capabilitie…

Reinforcement LearningInformation ExtractionText Spotting

Understanding Scanned Receipts

2020-05-04 · Eric Melz

Tasking machines with understanding receipts can have important applications such as enabling detailed analytics on purchases, enforcing expense policies, and inferring patterns of purchase behavior on large collections …

Entity LinkingInformation RetrievalRetrieval

Lack of corundum, carbon residues and revealing gaps on dental implants

2022-09-13 · Guy Florian Draenert, Gergo Mitov

Surface modification is an important topic to improve dental implants. Corundum residues, which are part of current dental implant blasting, disappeared on Straumann dental implants in recent publications. In our investi…