paper-with-me

Papers Document Image Classification

“Document Image Classification” 태그가 달린 논문 52편 · 필터 해제

Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models

2025-09-26 · Michael Jungo, Andreas Fischer arxiv

Rule-based reinforcement learning has been gaining popularity ever since DeepSeek-R1 has demonstrated its success through simple verifiable rewards. In the domain of document analysis, reinforcement learning is not as pr…

Document Image ClassificationReinforcement Learning

DocVCE: Diffusion-based Visual Counterfactual Explanations for Document Image Classification

2025-08-06 · Saifullah Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed arxiv

As black-box AI-driven decision-making systems become increasingly widespread in modern document processing workflows, improving their transparency and reliability has become critical, especially in high-stakes applicati…

Document Image ClassificationDocument Classification

Zero-Shot Prompting and Few-Shot Fine-Tuning: Revisiting Document Image Classification Using Large Language Models

2024-12-18 · Anna Scius-Bertrand, Michael Jungo, Lars Vögtlin, Jean-Marc Spat 외

Classifying scanned documents is a challenging problem that involves image, layout, and text analysis for document understanding. Nevertheless, for certain benchmark datasets, notably RVL-CDIP, the state of the art is cl…

Document Classificationdocument-image-classificationDocument Image Classificationdocument understanding+2

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

2024-12-17 · Nikitha SR, Tarun Ram Menta, Mausoom Sarkar

The advent of multimodal learning has brought a significant improvement in document AI. Documents are now treated as multimodal entities, incorporating both textual and visual information for downstream analysis. However…

Document AIDocument Image ClassificationDocument Layout AnalysisOptical Character Recognition (OCR)

DocXplain: A Novel Model-Agnostic Explainability Method for Document Image Classification

2024-07-04 · Saifullah Saifullah, Stefan Agne, Andreas Dengel, Sheraz Ahmed

Deep learning (DL) has revolutionized the field of document image analysis, showcasing superhuman performance across a diverse set of tasks. However, the inherent black-box nature of deep learning models still presents a…

document-image-classificationDocument Image ClassificationFairnessFeature Importance+2

DistilDoc: Knowledge Distillation for Visually-Rich Document Applications

2024-06-12 · Jordy Van Landeghem, Subhajit Maity, Ayan Banerjee, Matthew Blaschko 외

This work explores knowledge distillation (KD) for visually-rich document (VRD) applications such as document layout analysis (DLA) and document image classification (DIC). While VRD research is dependent on increasingly…

document-image-classificationDocument Image ClassificationDocument Layout Analysisdocument understanding+6

Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting

2024-05-21 · Omar Hamed, Souhail Bakkali, Marie-Francine Moens, Matthew Blaschko 외

This work addresses the need for a balanced approach between performance and efficiency in scalable production environments for visually-rich document understanding (VDU) tasks. Currently, there is a reliance on large do…

document-image-classificationDocument Image Classificationdocument understandingimage-classification+1

CICA: Content-Injected Contrastive Alignment for Zero-Shot Document Image Classification

2024-05-06 · Sankalp Sinha, Muhammad Saif Ullah Khan, Talha Uddin Sheikh, Didier Stricker 외

Zero-shot learning has been extensively investigated in the broader field of visual recognition, attracting significant interest recently. However, the current work on zero-shot learning in document image classification …

Document Classificationdocument-image-classificationDocument Image ClassificationGeneralized Zero-Shot Learning+3

LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding

2024-03-21 · Masato Fujitake

This paper proposes LayoutLLM, a more flexible document analysis method for understanding imaged documents. Visually Rich Document Understanding tasks, such as document image classification and information extraction, ha…

document-image-classificationDocument Image Classificationdocument understandingimage-classification+4

Automatic Recognition of Learning Resource Category in a Digital Library

2023-11-28 · Soumya Banerjee, Debarshi Kumar Sanyal, Samiran Chattopadhyay, Plaban Kumar Bhowmick 외

Digital libraries often face the challenge of processing a large volume of diverse document types. The manual collection and tagging of metadata can be a time-consuming and error-prone task. To address this, we aim to de…

document-image-classificationDocument Image Classificationimage-classificationImage Classification+1

SUT: a new multi-purpose synthetic dataset for Farsi document image analysis

2023-11-27 · 13th International Conference on Computer and Knowledge Engineering (ICCKE) 2023 11 · Elham Shabaninia, Fatemeh sadat Eslami, Ali Afkari Fahandari, Hossein Nezamabadi-pour

This paper introduces a new large-scale dataset for Farsi document images, named SUT, which aims to tackle the challenges associated with obtaining diverse and substantial ground-truth data for supervised models in docum…

Document Classificationdocument-image-classificationDocument Image Classificationimage-classification+5

A Multi-Modal Multilingual Benchmark for Document Image Classification

2023-10-25 · Yoshinari Fujinuma, Siddharth Varia, Nishant Sankaran, Srikar Appalaraju 외

Document image classification is different from plain-text document classification and consists of classifying a document by understanding the content and structure of documents such as forms, emails, and other such docu…

ClassificationCross-Lingual TransferDocument AIDocument Classification+8

GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification

2023-09-11 · Souhail Bakkali, Sanket Biswas, Zuheng Ming, Mickaël Coustaty 외

Visual document understanding (VDU) has rapidly advanced with the development of powerful multi-modal language models. However, these models typically require extensive document pre-training data to learn intermediate re…

document-image-classificationDocument Image Classificationdocument understandingimage-classification+3

LayoutMask: Enhance Text-Layout Interaction in Multi-modal Pre-training for Document Understanding

2023-05-30 · Yi Tu, Ya Guo, Huan Chen, Jinyang Tang

Visually-rich Document Understanding (VrDU) has attracted much research attention over the past years. Pre-trained models on a large number of document images with transformer-based backbones have led to significant perf…

document-image-classificationDocument Image Classificationdocument understandingimage-classification+9

EAML: Ensemble Self-Attention-based Mutual Learning Network for Document Image Classification

2023-05-11 · IJDAR 2021 6 · Souhail Bakkali, Ziheng Ming, Mickael Coustaty, Marçal Rusiñol

In the recent past, complex deep neural networks have received huge interest in various document understanding tasks such as document image classification and document retrieval. As many document types have a distinct vi…

document-image-classificationDocument Image Classificationimage-classification

Evaluating Adversarial Robustness on Document Image Classification

2023-04-24 · Timothée Fronteau, Arnaud Paran, Aymen Shabou

Adversarial attacks and defenses have gained increasing interest on computer vision systems in recent years, but as of today, most investigations are limited to images. However, many artificial intelligence models actual…

Adversarial AttackAdversarial RobustnessClassificationdocument-image-classification+4

Context-Aware Classification of Legal Document Pages

2023-04-05 · Pavlos Fragkogiannis, Martina Forster, Grace E. Lee, Dell Zhang

For many business applications that require the processing, indexing, and retrieval of professional documents such as legal briefs (in PDF format etc.), it is often essential to classify the pages of any given document i…

Classificationdocument-image-classificationDocument Image Classificationimage-classification+2

StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

2023-03-01 · Yuechen Yu, Yulin Li, Chengquan Zhang, Xiaoqiang Zhang 외

In this paper, we present StrucTexTv2, an effective document image pre-training framework, by performing masked visual-textual prediction. It consists of two self-supervised pre-training tasks: masked image modeling and …

Document Image Classificationimage-classificationImage ClassificationLanguage Modeling+4

Multimodal Side-Tuning for Document Classification

2023-01-16 · Stefano Pio Zingaro, Giuseppe Lisanti, Maurizio Gabbrielli

In this paper, we propose to exploit the side-tuning framework for multimodal document classification. Side-tuning is a methodology for network adaptation recently introduced to solve some of the problems related to prev…

ClassificationDocument ClassificationDocument Image ClassificationTransfer Learning

ERNIE-Layout: Layout Knowledge Enhanced Pre-training for Visually-rich Document Understanding

2022-10-12 · Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo 외

Recent years have witnessed the rise and success of pre-training techniques in visually-rich document understanding. However, most existing methods lack the systematic mining and utilization of layout-centered knowledge,…

document-image-classificationDocument Image Classificationdocument understandingimage-classification+6
1–20 / 52 다음 →