paper-with-me

Papers

AMuRD: Annotated Arabic-English Receipt Dataset for Key Information Extraction and Classification

2023-09-18 · Abdelrahman Abdallah, Mahmoud Abdalla, Mohamed Elkasaby, Yasser Elbendary, Adam Jatowt

The extraction of key information from receipts is a complex task that involves the recognition and extraction of text from scanned receipts. This process is crucial as it enables the retrieval of essential content and organizing it into structured documents for easy access and analysis. In this paper, we present AMuRD, a novel multilingual human-annotated dataset specifically designed for information extraction from receipts. This dataset comprises $47,720$ samples and addresses the key challenges in information extraction and item classification - the two critical aspects of data analysis in the retail industry. Each sample includes annotations for item names and attributes such as price, brand, and more. This detailed annotation facilitates a comprehensive understanding of each item on the receipt. Furthermore, the dataset provides classification into $44$ distinct product categories. This classification feature allows for a more organized and efficient analysis of the items, enhancing the usability of the dataset for various applications. In our study, we evaluated various language model architectures, e.g., by fine-tuning LLaMA models on the AMuRD dataset. Our approach yielded exceptional results, with an F1 score of 97.43\% and accuracy of 94.99\% in information extraction and classification, and an even higher F1 score of 98.51\% and accuracy of 97.06\% observed in specific tasks. The dataset and code are publicly accessible for further researchhttps://github.com/Update-For-Integrated-Business-AI/AMuRD.

📄 PDF Abstract BibTeX arXiv:2309.09800

Code (1)

update-for-integrated-business-ai/amurd 공식 구현 pytorch

Tasks

ClassificationKey Information ExtractionLanguage ModellingRetrieval

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

CORU: Comprehensive Post-OCR Parsing and Receipt Understanding Dataset

2024-06-06 · Abdelrahman Abdallah, Mahmoud Abdalla, Mahmoud SalahEldin Kasem, Mohamed Mahmoud 외

In the fields of Optical Character Recognition (OCR) and Natural Language Processing (NLP), integrating multilingual capabilities remains a critical challenge, especially when considering languages with complex scripts s…

object-detectionObject DetectionOptical Character RecognitionOptical Character Recognition (OCR)

EmoHopeSpeech: An Annotated Dataset of Emotions and Hope Speech in English and Arabic

2025-05-17 · Wajdi Zaghouani, Md. Rafiul Biswas

This research introduces a bilingual dataset comprising 23,456 entries for Arabic and 10,036 entries for English, annotated for emotions and hope speech, addressing the scarcity of multi-emotion (Emotion and hope) datase…

Morphologically Annotated Corpora for Seven Arabic Dialects: Taizi, Sanaani, Najdi, Jordanian, Syrian, Iraqi and Moroccan

2019-08-01 · WS 2019 8 · Faisal Alshargi, Shahd Dibas, Sakhar Alkhereyf, Reem Faraj 외

We present a collection of morphologically annotated corpora for seven Arabic dialects: Taizi Yemeni, Sanaani Yemeni, Najdi, Jordanian, Syrian, Iraqi and Moroccan Arabic. The corpora collectively cover over 200,000 words…

Morphological Analysis

Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation

2025-08-27 · Slimane Bellaouar, Attia Nehar, Soumia Souffi, Mounia Bouameur arxiv

Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces the challenge of being under-resourced. The scarcity of large annotated datasets hampers the development of accurate too…

Subjectivity AnalysisText Classification

CARMA: Comprehensive Automatically-annotated Reddit Mental Health Dataset for Arabic

2025-11-05 · Saad Mankarious, Ayah Zirikly arxiv

Mental health disorders affect millions worldwide, yet early detection remains a major challenge, particularly for Arabic-speaking populations where resources are limited and mental health discourse is often discouraged …