paper-with-me

Document Image Classification

8개 벤치마크 · 논문 52편 · 이 태스크의 논문 보기 →

Benchmarks

RVL-CDIP

결과 62개

Tobacco-3482

결과 20개

Noisy Bangla Numeral

결과 4개

AIP

결과 2개

Noisy MNIST

결과 2개

SUT

결과 2개

n-MNIST

결과 2개

Most implemented

Papers

Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models

2025-09-26 · Michael Jungo, Andreas Fischer arxiv

Rule-based reinforcement learning has been gaining popularity ever since DeepSeek-R1 has demonstrated its success through simple verifiable rewards. In the domain of document analysis, reinforcement learning is not as pr…

Document Image ClassificationReinforcement Learning

DocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification

2025-08-06 · Saifullah Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed arxiv

As black-box AI-driven decision-making systems become increasingly widespread in modern document processing workflows, improving their transparency and reliability has become critical, especially in high-stakes applicati…

Document Image ClassificationDocument Classification

Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models

2024-12-18 · Anna Scius-Bertrand, Michael Jungo, Lars Vögtlin, Jean-Marc Spat 외

Classifying scanned documents is a challenging problem that involves image, layout, and text analysis for document understanding. Nevertheless, for certain benchmark datasets, notably RVL-CDIP, the state of the art is cl…

Document Classificationdocument-image-classificationDocument Image Classificationdocument understanding+2

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

2024-12-17 · Nikitha SR, Tarun Ram Menta, Mausoom Sarkar

The advent of multimodal learning has brought a significant improvement in document AI. Documents are now treated as multimodal entities, incorporating both textual and visual information for downstream analysis. However…

Document AIDocument Image ClassificationDocument Layout AnalysisOptical Character Recognition (OCR)

DocXplain: A Novel Model-Agnostic Explainability Method for Document Image Classification

2024-07-04 · Saifullah Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed

Deep learning (DL) has revolutionized the field of document image analysis, showcasing superhuman performance across a diverse set of tasks. However, the inherent black-box nature of deep learning models still presents a…

document-image-classificationDocument Image ClassificationFairnessFeature Importance+2

DistilDoc: Knowledge Distillation for Visually-Rich Document Applications

2024-06-12 · Jordy Van Landeghem, Subhajit Maity, Ayan Banerjee, Matthew Blaschko 외

This work explores knowledge distillation (KD) for visually-rich document (VRD) applications such as document layout analysis (DLA) and document image classification (DIC). While VRD research is dependent on increasingly…

document-image-classificationDocument Image ClassificationDocument Layout Analysisdocument understanding+6

전체 52편 보기 →