paper-with-me

홈 › Papers

Multimodal deep networks for text and image-based document classification

2019-07-15 · Nicolas Audebert, Catherine Herold, Kuider Slimani, Cédric Vidal

Classification of document images is a critical step for archival of old manuscripts, online subscription and administrative procedures. Computer vision and deep learning have been suggested as a first solution to classify documents based on their visual appearance. However, achieving the fine-grained classification that is required in real-world setting cannot be achieved by visual analysis alone. Often, the relevant information is in the actual text content of the document. We design a multimodal neural network that is able to learn from word embeddings, computed on text extracted by OCR, and from the image. We show that this approach boosts pure image accuracy by 3% on Tobacco3482 and RVL-CDIP augmented by our new QS-OCR text dataset (https://github.com/Quicksign/ocrized-text-dataset), even without clean text information.

📄 PDF Abstract BibTeX arXiv:1907.06370

Code (3)

Quicksign/ocrized-text-dataset 공식 구현
bharathrajcl/Multimodal-deep-networks-for-text-and-image-based-document-classification tf
mleimeister/document-image-classification tf

Tasks

ClassificationDocument ClassificationGeneral ClassificationMultimodal Deep LearningOptical Character Recognition (OCR)Word Embeddings

Similar Papers 제목 키워드 기반

Deep Learning for Technical Document Classification

2021-06-27 · Shuo Jiang, Jie Hu, Christopher L. Magee, Jianxi Luo

In large technology companies, the requirements for managing and organizing technical documents created by engineers and managers have increased dramatically in recent years, which has led to a higher demand for more sca…

ClassificationDecision MakingDeep LearningDescriptive+6

Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis

2026-06-01 · Catyana Heyne, Jürgen Frikel, Filippo Riccio arxiv

Document type classification in visually rich documents remains challenging, as relevant information is distributed across textual, visual, and layout modalities. To capture this complexity, current approaches rely on di…

Visual Word Embedding for Text Classification

2021-02-25 · Ignazio Gallo, Shah Nawaz, Nicola Landro, Riccardo La Grassa

The question we answer with this paper is: ‘can we convert a text document into an image to take advantage of image neural models to classify text documents?’ To answer this question we present a novel text classificatio…

ClassificationGeneral Classificationimage-classificationImage Classification+2

LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

2022-04-18 · Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu 외

Self-supervised pre-training techniques have achieved remarkable progress in Document AI. Most multimodal pre-trained models use a masked language modeling objective to learn bidirectional representations on the text mod…

cross-modal alignmentDocument AIdocument-image-classificationDocument Image Classification+16

DocXClassifier: High Performance Explainable Deep Network for Document Image Classification

2022-03-17 · TechArXiv 2022 3 · Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed

Convolutional Neural Networks (ConvNets) have been thoroughly researched for document image classification and are known for their exceptional performance in unimodal image-based document classification. Recently, howe…

ClassificationData AugmentationDocument Classificationdocument-image-classification+5