paper-with-me

홈 › Papers

Visual and Textual Deep Feature Fusion for Document Image Classification

2020-06-16 · CVPRW 2020 6 · Souhail Bakkali, Ziheng Ming, Mickael Coustaty, Marçal Rusiñol

The topic of text document image classification has been explored extensively over the past few years. Most recent approaches handled this task by jointly learning the visual features of document images and their corresponding textual contents. Due to the various structures of document images, the extraction of semantic information from its textual content is beneficial for document image processing tasks such as document retrieval, information extraction, and text classification. In this work, a two-stream neural architecture is proposed to perform the document image classification task. We conduct an exhaustive investigation of nowadays widely used neural networks as well as word embedding procedures used as backbones, in order to extract both visual and textual features from document images. Moreover, a joint feature learning approach that combines image features and text embeddings is introduced as a late fusion methodology. Both the theoretical analysis and the experimental results demonstrate the superiority of our proposed joint feature learning method comparatively to the single modalities. This joint learning approach outperforms the state-of-the-art results with a classification accuracy of 97.05% on the large-scale RVL-CDIP dataset.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Classificationdocument-image-classificationDocument Image Classificationimage-classificationImage ClassificationRetrievaltext-classificationText Classification

Similar Papers 제목 키워드 기반

SelfDoc: Self-Supervised Document Representation Learning

2021-06-07 · CVPR 2021 1 · Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu 외

We propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework exploits the positional, textual, and v…

Representation Learning

Sequence-aware multimodal page classification of Brazilian legal documents

2022-07-02 · Pedro H. Luz de Araujo, Ana Paula G. S. de Almeida, Fabricio A. Braz, Nilton C. da Silva 외

The Brazilian Supreme Court receives tens of thousands of cases each semester. Court employees spend thousands of hours to execute the initial analysis and classification of those cases -- which takes effort away from po…

ClassificationManagementOptical Character RecognitionOptical Character Recognition (OCR)

Multimodal Pre-training Based on Graph Attention Network for Document Understanding

2022-03-25 · Zhenrong Zhang, Jiefeng Ma, Jun Du, Licheng Wang 외

Document intelligence as a relatively new research topic supports many business applications. Its main task is to automatically read, understand, and analyze documents. However, due to the diversity of formats (invoices,…

document understandingGraph AttentionSentence

MiMIC: Mitigating Visual Modality Collapse in Universal Multimodal Retrieval While Avoiding Semantic Misalignment

2026-04-23 · Juan Li, Chuanghao Ding, Xujie Zhang, Cam-Tu Nguyen arxiv

Universal Multimodal Retrieval (UMR) aims to map different modalities (e.g., visual and textual) into a shared embedding space for multi-modal retrieval. Existing UMR methods can be broadly divided into two categories: e…

Combining Visual and Textual Features for Semantic Segmentation of Historical Newspapers

2020-02-14 · Raphaël Barman, Maud Ehrmann, Simon Clematide, Sofia Ares Oliveira 외

The massive amounts of digitized historical documents acquired over the last decades naturally lend themselves to automatic processing and exploration. Research work seeking to automatically process facsimiles and extrac…

Document Layout AnalysisSemantic Segmentation