paper-with-me

홈 › Papers

Test-Time Adaptation for Visual Document Understanding

2022-06-15 · Sayna Ebrahimi, Sercan O. Arik, Tomas Pfister

For visual document understanding (VDU), self-supervised pretraining has been shown to successfully generate transferable representations, yet, effective adaptation of such representations to distribution shifts at test-time remains to be an unexplored area. We propose DocTTA, a novel test-time adaptation method for documents, that does source-free domain adaptation using unlabeled target document data. DocTTA leverages cross-modality self-supervised learning via masked visual language modeling, as well as pseudo labeling to adapt models learned on a \textit{source} domain to an unlabeled \textit{target} domain at test time. We introduce new benchmarks using existing public datasets for various VDU tasks, including entity recognition, key-value extraction, and document visual question answering. DocTTA shows significant improvements on these compared to the source model performance, up to 1.89\% in (F1 score), 3.43\% (F1 score), and 17.68\% (ANLS score), respectively. Our benchmark datasets are available at \url{https://saynaebrahimi.github.io/DocTTA.html}.

📄 PDF Abstract BibTeX arXiv:2206.07240

Code (0)

등록된 구현이 없습니다.

Tasks

document understandingDomain AdaptationLanguage ModelingLanguage ModellingQuestion AnsweringSelf-Supervised LearningSource-Free Domain AdaptationTest-time AdaptationVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Visual Document Understanding and Reasoning: A Multi-Agent Collaboration Framework with Agent-Wise Adaptive Test-Time Scaling

2025-08-05 · Xinlei Yu, Chengming Xu, Zhangquan Chen, Yudong Zhang 외 arxiv

The dominant paradigm of monolithic scaling in Vision-Language Models (VLMs) is failing for understanding and reasoning in documents, yielding diminishing returns as it struggles with the inherent need of this domain for…

Mathematical Reasoning

LoRA-Contextualizing Adaptation of Large Multimodal Models for Long Document Understanding

2024-11-02 · Jian Chen, Ruiyi Zhang, Yufan Zhou, Tong Yu 외

Large multimodal models (LMMs) have recently shown great progress in text-rich image understanding, yet they still struggle with complex, multi-page, visually-rich documents. Traditional methods using document parsers fo…

document understandingQuestion AnsweringRetrievalRetrieval-augmented Generation

DAViD: Domain Adaptive Visually-Rich Document Understanding with Synthetic Insights

2024-10-02 · Yihao Ding, Soyeon Caren Han, Zechuan Li, Hyunsuk Chung

Visually-Rich Documents (VRDs), encompassing elements like charts, tables, and references, convey complex information across various fields. However, extracting information from these rich documents is labor-intensive, e…

document understandingDomain AdaptationRepresentation Learning

MonkeyOCRv2: A Visual-Text Foundation Model for Document AI

2026-07-13 · Yuliang Liu, Zhang Li, Ziyang Zhang, Shuo Zhang 외 hf

Mainstream visual encoders are pretrained on natural images and cannot be effectively applied to document images without document-oriented adaptation, as dense text and fine-grained character strokes demand character-lev…

Text GenerationText DetectionDocument AI

SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction

2025-09-27 · Yihao Ding, Soyeon Caren Han, Yanbei Jiang, Yan Li 외 arxiv

Domain-specific Visually Rich Document Understanding (VRDU) presents significant challenges due to the complexity and sensitivity of documents in fields such as medicine, finance, and material science. Existing Large (Mu…

Key Information ExtractionSynthetic Data GenerationDomain Adaptation