paper-with-me

Papers

Modular Multimodal Architecture for Document Classification

2019-12-09 · Tyler Dauphinee, Nikunj Patel, Mohammad Rashidi

Page classification is a crucial component to any document analysis system, allowing for complex branching control flows for different components of a given document. Utilizing both the visual and textual content of a page, the proposed method exceeds the current state-of-the-art performance on the RVL-CDIP benchmark at 93.03% test accuracy.

📄 PDF Abstract BibTeX arXiv:1912.04376

Code (3)

BordiaS/layoutlm pytorch
iamarjunchandra/LayoutLM-Form-Understanding---Sequence-Labeling pytorch
microsoft/unilm/tree/master/layoutlm pytorch

Tasks

ClassificationDocument ClassificationGeneral Classification

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Modular Multimodal Machine Learning for Extraction of Theorems and Proofs in Long Scientific Documents (Extended Version)

2023-07-18 · Shrey Mishra, Antoine Gauquier, Pierre Senellart

We address the extraction of mathematical statements and their proofs from scholarly PDF articles as a multimodal classification problem, utilizing text, font features, and bitmap image renderings of PDFs as distinct mod…

ArticlesDocument AILanguage ModellingOptical Character Recognition (OCR)

Deep Learning for Technical Document Classification

2021-06-27 · Shuo Jiang, Jie Hu, Christopher L. Magee, Jianxi Luo

In large technology companies, the requirements for managing and organizing technical documents created by engineers and managers have increased dramatically in recent years, which has led to a higher demand for more sca…

ClassificationDecision MakingDeep LearningDescriptive+6

PostDoc: Generating Poster from a Long Multimodal Document Using Deep Submodular Optimization

2024-05-30 · Vijay Jaisankar, Sambaran Bandyopadhyay, Kalp Vyas, Varre Chaitanya 외

A poster from a long input document can be considered as a one-page easy-to-read multimodal (text and images) summary presented on a nice template with good design elements. Automatic transformation of a long document in…

Diversity

Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis

2026-06-01 · Catyana Heyne, Jürgen Frikel, Filippo Riccio arxiv

Document type classification in visually rich documents remains challenging, as relevant information is distributed across textual, visual, and layout modalities. To capture this complexity, current approaches rely on di…

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval

2026-01-31 · Kiarash Naghavi Khanghah, Hoang Anh Nguyen, Anna C. Doris, Amir Mohammad Vahedi 외 arxiv

Engineering rulebooks and technical standards contain multimodal information like dense text, tables, and illustrations that are challenging for retrieval augmented generation (RAG) systems. Building upon the DesignQA fr…

Question Answering