Papers Document Image Classification
“Document Image Classification” 태그가 달린 논문 52편 · 필터 해제
Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
Rule-based reinforcement learning has been gaining popularity ever since DeepSeek-R1 has demonstrated its success through simple verifiable rewards. In the domain of document analysis, reinforcement learning is not as pr…
Document Image ClassificationReinforcement LearningDocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification
As black-box AI-driven decision-making systems become increasingly widespread in modern document processing workflows, improving their transparency and reliability has become critical, especially in high-stakes applicati…
Document Image ClassificationDocument ClassificationZero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models
Classifying scanned documents is a challenging problem that involves image, layout, and text analysis for document understanding. Nevertheless, for certain benchmark datasets, notably RVL-CDIP, the state of the art is cl…
Document Classificationdocument-image-classificationDocument Image Classificationdocument understanding+2DoPTA: Improving Document Layout Analysis using Patch-Text Alignment
The advent of multimodal learning has brought a significant improvement in document AI. Documents are now treated as multimodal entities, incorporating both textual and visual information for downstream analysis. However…
Document AIDocument Image ClassificationDocument Layout AnalysisOptical Character Recognition (OCR)DocXplain: A Novel Model-Agnostic Explainability Method for Document Image Classification
Deep learning (DL) has revolutionized the field of document image analysis, showcasing superhuman performance across a diverse set of tasks. However, the inherent black-box nature of deep learning models still presents a…
document-image-classificationDocument Image ClassificationFairnessFeature Importance+2DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
This work explores knowledge distillation (KD) for visually-rich document (VRD) applications such as document layout analysis (DLA) and document image classification (DIC). While VRD research is dependent on increasingly…
document-image-classificationDocument Image ClassificationDocument Layout Analysisdocument understanding+6Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
This work addresses the need for a balanced approach between performance and efficiency in scalable production environments for visually-rich document understanding (VDU) tasks. Currently, there is a reliance on large do…
document-image-classificationDocument Image Classificationdocument understandingimage-classification+1CICA: Content-Injected Contrastive Alignment for Zero-Shot Document Image Classification
Zero-shot learning has been extensively investigated in the broader field of visual recognition, attracting significant interest recently. However, the current work on zero-shot learning in document image classification …
Document Classificationdocument-image-classificationDocument Image ClassificationGeneralized Zero-Shot Learning+3LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
This paper proposes LayoutLLM, a more flexible document analysis method for understanding imaged documents. Visually Rich Document Understanding tasks, such as document image classification and information extraction, ha…
document-image-classificationDocument Image Classificationdocument understandingimage-classification+4Automatic Recognition of Learning Resource Category in a Digital Library
Digital libraries often face the challenge of processing a large volume of diverse document types. The manual collection and tagging of metadata can be a time-consuming and error-prone task. To address this, we aim to de…
document-image-classificationDocument Image Classificationimage-classificationImage Classification+1SUT: a new multi-purpose synthetic dataset for Farsi document image analysis
This paper introduces a new large-scale dataset for Farsi document images, named SUT, which aims to tackle the challenges associated with obtaining diverse and substantial ground-truth data for supervised models in docum…
Document Classificationdocument-image-classificationDocument Image Classificationimage-classification+5A Multi-Modal Multilingual Benchmark for Document Image Classification
Document image classification is different from plain-text document classification and consists of classifying a document by understanding the content and structure of documents such as forms, emails, and other such docu…
ClassificationCross-Lingual TransferDocument AIDocument Classification+8GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
Visual document understanding (VDU) has rapidly advanced with the development of powerful multi-modal language models. However, these models typically require extensive document pre-training data to learn intermediate re…
document-image-classificationDocument Image Classificationdocument understandingimage-classification+3LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding
Visually-rich Document Understanding (VrDU) has attracted much research attention over the past years. Pre-trained models on a large number of document images with transformer-based backbones have led to significant perf…
document-image-classificationDocument Image Classificationdocument understandingimage-classification+9EAML: Ensemble Self-Attention-based Mutual Learning Network for Document Image Classification
In the recent past, complex deep neural networks have received huge interest in various document understanding tasks such as document image classification and document retrieval. As many document types have a distinct vi…
document-image-classificationDocument Image Classificationimage-classificationEvaluating Adversarial Robustness on Document Image Classification
Adversarial attacks and defenses have gained increasing interest on computer vision systems in recent years, but as of today, most investigations are limited to images. However, many artificial intelligence models actual…
Adversarial AttackAdversarial RobustnessClassificationdocument-image-classification+4Context-Aware Classification of Legal Document Pages
For many business applications that require the processing, indexing, and retrieval of professional documents such as legal briefs (in PDF format etc.), it is often essential to classify the pages of any given document i…
Classificationdocument-image-classificationDocument Image Classificationimage-classification+2StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training
In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-training tasks: masked image modeling and …
Document Image Classificationimage-classificationImage ClassificationLanguage Modeling+4Multimodal Side-Tuning for Document Classification
In this paper, we propose to exploit the side-tuning framework for multimodal document classification. Side-tuning is a methodology for network adaptation recently introduced to solve some of the problems related to prev…
ClassificationDocument ClassificationDocument Image ClassificationTransfer LearningERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding
Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and utilization of layout-centered knowledge,…
document-image-classificationDocument Image Classificationdocument understandingimage-classification+6