paper-with-me

홈 › Papers

DocXClassifier: High Performance Explainable Deep Network for Document Image Classification

2022-03-17 · TechArXiv 2022 3 · Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed

Convolutional Neural Networks (ConvNets) have been thoroughly researched for document image classification and are known for their exceptional performance in unimodal image-based document classification. Recently, however, there has been a sudden shift in the field towards multimodal approaches that simultaneously learn from the visual and textual features of the documents. While this has led to significant advances in the field, it has also led to a waning interest in improving pure ConvNets-based approaches. This is not desirable, as many of the multimodal approaches still use ConvNets as their visual backbone, and thus improving ConvNets is essential to improving these approaches. In this paper, we present DocXClassifier, a ConvNet-based approach that, using state-of-the-art model design patterns together with modern data augmentation and training strategies, not only achieves significant performance improvements in image-based document classification, but also outperforms some of the recently proposed multimodal approaches. Moreover, DocXClassifier is capable of generating transformer-like attention maps, which makes it inherently interpretable, a property not found in previous image-based classification models. Our approach achieves a new peak performance in image-based classification on two popular document datasets, namely RVL-CDIP and Tobacco3482, with a top-1 classification accuracy of 94.17% and 95.57% on the two datasets, respectively. Moreover, it sets a new record for the highest image-based classification accuracy of 90.14% on Tobacco3482 without transfer learning from RVL-CDIP. Finally, our proposed model may serve as a powerful visual backbone for future multimodal approaches, by providing much richer visual features than existing counterparts.

📄 PDF Abstract BibTeX

Code (1)

saifullah3396/docxclassifier pytorch

Tasks

ClassificationData AugmentationDocument Classificationdocument-image-classificationDocument Image Classificationimage-classificationImage ClassificationTransfer LearningVocal Bursts Intensity Prediction

Similar Papers 제목 키워드 기반

WordVIS: A Color Worth A Thousand Words

2024-12-13 · Umar Khan, Saifullah, Stefan Agne, Andreas Dengel 외

Document classification is considered a critical element in automated document processing systems. In recent years multi-modal approaches have become increasingly popular for document classification. Despite their improv…

Document Classification

Explainable Text Classification in Legal Document Review A Case Study of Explainable Predictive Coding

2019-04-03 · Rishi Chhatwal, Peter Gronvall, Nathaniel Huber-Fliflet, Robert Keeling 외

In today's legal environment, lawsuits and regulatory investigations require companies to embark upon increasingly intensive data-focused engagements to identify, collect and analyze large quantities of data. When docume…

Document ClassificationGeneral Classificationtext-classificationText Classification

Explainable identification of similarities between entities for discovery in large text

2025-03-22 · Akhil Joshi, Sai Teja Erukude, Lior Shamir

With the availability of virtually infinite number text documents in digital format, automatic comparison of textual data is essential for extracting meaningful insights that are difficult to identify manually. Many exis…

Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions

2026-05-07 · Kjetil Indrehus, Adrian Duric, Changkyu Choi, Ali Ramezani-Kebrya arxiv

Document Visual Question Answering (DocVQA) requires vision-language models to reason not only about what information in a document is relevant to a question, but also where the answer is grounded on the page. Existing D…

Visual Question Answering

A Framework for Explainable Text Classification in Legal Document Review

2019-12-19 · Christian J. Mahoney, Jianping Zhang, Nathaniel Huber-Fliflet, Peter Gronvall 외

Companies regularly spend millions of dollars producing electronically-stored documents in legal matters. Recently, parties on both sides of the 'legal aisle' are accepting the use of machine learning techniques like tex…

ClassificationGeneral Classificationtext-classificationText Classification