paper-with-me

Papers

AraDIC: Arabic Document Classification using Image-Based Character Embeddings and Class-Balanced Loss

2020-06-20 · ACL 2020 6 · Mahmoud Daif, Shunsuke Kitada, Hitoshi Iyatomi

Classical and some deep learning techniques for Arabic text classification often depend on complex morphological analysis, word segmentation, and hand-crafted feature engineering. These could be eliminated by using character-level features. We propose a novel end-to-end Arabic document classification framework, Arabic document image-based classifier (AraDIC), inspired by the work on image-based character embeddings. AraDIC consists of an image-based character encoder and a classifier. They are trained in an end-to-end fashion using the class balanced loss to deal with the long-tailed data distribution problem. To evaluate the effectiveness of AraDIC, we created and published two datasets, the Arabic Wikipedia title (AWT) dataset and the Arabic poetry (AraP) dataset. To the best of our knowledge, this is the first image-based character embedding framework addressing the problem of Arabic text classification. We also present the first deep learning-based text classifier widely evaluated on modern standard Arabic, colloquial Arabic and classical Arabic. AraDIC shows performance improvement over classical and deep learning baselines by 12.29% and 23.05% for the micro and macro F-score, respectively.

📄 PDF Abstract BibTeX arXiv:2006.11586

Code (1)

mahmouddaif/AraDIC 공식 구현 pytorch

Tasks

ClassificationDeep LearningDocument ClassificationFeature EngineeringGeneral ClassificationMorphological Analysistext-classificationText Classification

Similar Papers 제목 키워드 기반

AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs

2024-09-17 · Basel Mousi, Nadir Durrani, Fatema Ahmad, Md. Arid Hasan 외

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects …

Dialect IdentificationDiversityMachine TranslationTranslation

An Efficient Language-Independent Multi-Font OCR for Arabic Script

2020-09-18 · Hussein Osman, Karim Zaghw, Mostafa Hazem, Seifeldin Elsehely

Optical Character Recognition (OCR) is the process of extracting digitized text from images of scanned documents. While OCR systems have already matured in many languages, they still have shortcomings in cursive language…

Optical Character RecognitionOptical Character Recognition (OCR)Segmentation

Compilation of an Arabic Children's Corpus

2016-05-01 · LREC 2016 5 · Latifa Al-Sulaiti, Noorhan Abbas, Claire Brierley, Eric Atwell 외

Inspired by the Oxford Children{'}s Corpus, we have developed a prototype corpus of Arabic texts written and/or selected for children. Our Arabic Children{'}s Corpus of 2950 documents and nearly 2 million words has been …

General Classificationtext-classificationText Classification

Arabic Character Segmentation Using Projection Based Approach with Profile's Amplitude Filter

2017-07-04 · Mahmoud A. A. Mousa, Mohammed S. Sayed, Mahmoud I. Abdalla

Arabic is one of the languages that present special challenges to Optical character recognition (OCR). The main challenge in Arabic is that it is mostly cursive. Therefore, a segmentation process must be carried out to d…

Optical Character RecognitionOptical Character Recognition (OCR)Segmentation

Cross-Lingual SynthDocs: A Large-Scale Synthetic Corpus for Any to Arabic OCR and Document Understanding

2025-11-01 · Haneen Al-Homoud, Asma Ibrahim, Murtadha Al-Jubran, Fahad Al-Otaibi 외 arxiv

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 milli…