Document Image Classification
8개 벤치마크 · 논문 52편 · 이 태스크의 논문 보기 →
Benchmarks
RVL-CDIP
Tobacco-3482
Noisy Bangla Characters
Noisy Bangla Numeral
AIP
Noisy MNIST
SUT
n-MNIST
Most implemented
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Training data-efficient image transformers & distillation through attention
LayoutLM: Pre-training of Text and Layout for Document Image Understanding
BEiT: BERT Pre-Training of Image Transformers
Papers
Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
Rule-based reinforcement learning has been gaining popularity ever since DeepSeek-R1 has demonstrated its success through simple verifiable rewards. In the domain of document analysis, reinforcement learning is not as pr…
Document Image ClassificationReinforcement LearningDocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification
As black-box AI-driven decision-making systems become increasingly widespread in modern document processing workflows, improving their transparency and reliability has become critical, especially in high-stakes applicati…
Document Image ClassificationDocument ClassificationZero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models
Classifying scanned documents is a challenging problem that involves image, layout, and text analysis for document understanding. Nevertheless, for certain benchmark datasets, notably RVL-CDIP, the state of the art is cl…
Document Classificationdocument-image-classificationDocument Image Classificationdocument understanding+2DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
The advent of multimodal learning has brought a significant improvement in document AI. Documents are now treated as multimodal entities, incorporating both textual and visual information for downstream analysis. However…
Document AIDocument Image ClassificationDocument Layout AnalysisOptical Character Recognition (OCR)DocXplain: A Novel Model-Agnostic Explainability Method for Document Image Classification
Deep learning (DL) has revolutionized the field of document image analysis, showcasing superhuman performance across a diverse set of tasks. However, the inherent black-box nature of deep learning models still presents a…
document-image-classificationDocument Image ClassificationFairnessFeature Importance+2DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
This work explores knowledge distillation (KD) for visually-rich document (VRD) applications such as document layout analysis (DLA) and document image classification (DIC). While VRD research is dependent on increasingly…
document-image-classificationDocument Image ClassificationDocument Layout Analysisdocument understanding+6