paper-with-me

홈 › Papers

Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism

2024-04-29 · Lei Kang, Rubèn Tito, Ernest Valveny, Dimosthenis Karatzas

Documents are 2-dimensional carriers of written communication, and as such their interpretation requires a multi-modal approach where textual and visual information are efficiently combined. Document Visual Question Answering (Document VQA), due to this multi-modal nature, has garnered significant interest from both the document understanding and natural language processing communities. The state-of-the-art single-page Document VQA methods show impressive performance, yet in multi-page scenarios, these methods struggle. They have to concatenate all pages into one large page for processing, demanding substantial GPU resources, even for evaluation. In this work, we propose a novel method and efficient training strategy for multi-page Document VQA tasks. In particular, we employ a visual-only document representation, leveraging the encoder from a document understanding model, Pix2Struct. Our approach utilizes a self-attention scoring mechanism to generate relevance scores for each document page, enabling the retrieval of pertinent pages. This adaptation allows us to extend single-page Document VQA models to multi-page scenarios without constraints on the number of pages during evaluation, all with minimal demand for GPU resources. Our extensive experiments demonstrate not only achieving state-of-the-art performance without the need for Optical Character Recognition (OCR), but also sustained performance in scenarios extending to documents of nearly 800 pages compared to a maximum of 20 pages in the MP-DocVQA dataset. Our code is publicly available at \url{https://github.com/leitro/SelfAttnScoring-MPDocVQA}.

📄 PDF Abstract BibTeX arXiv:2404.19024

Code (1)

leitro/selfattnscoring-mpdocvqa 공식 구현 pytorch

Tasks

document understandingGPUOptical Character RecognitionOptical Character Recognition (OCR)Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering

2026-01-17 · Zongmin Li, Yachuan Li, Lei Kang, Dimosthenis Karatzas 외 arxiv

Multi-page Document Visual Question Answering (MP-DocVQA) remains challenging because long documents not only strain computational resources but also reduce the effectiveness of the attention mechanism in large vision-la…

Visual Question AnsweringAnswer Generation

Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents

2025-11-14 · Davide Napolitano, Luca Cagliero, Fabrizio Battiloro arxiv

The evolution of Visual Large Language Models (VLLMs) has revolutionized the automatic understanding of Visually Rich Documents (VRDs), which contain both textual and visual elements. Although VLLMs excel in Visual Quest…

Visual Question Answering

Hierarchical multimodal transformers for Multi-Page DocVQA

2022-12-07 · Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny

Document Visual Question Answering (DocVQA) refers to the task of answering questions from document images. Existing work on DocVQA only considers single-page documents. However, in real scenarios documents are mostly co…

DecoderQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

2024-09-05 · Anwen Hu, Haiyang Xu, Liang Zhang, Jiabo Ye 외

Multimodel Large Language Models(MLLMs) have achieved promising OCR-free Document Understanding performance by increasing the supported resolution of document images. However, this comes at the cost of generating thousan…

document understandingGPUOptical Character Recognition (OCR)Question Answering

PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering

2024-04-19 · Yihao Ding, Kaixuan Ren, Jiabin Huang, Siwen Luo 외

Document Question Answering (QA) presents a challenge in understanding visually-rich documents (VRD), particularly those dominated by lengthy textual content like research journal articles. Existing studies primarily foc…

ArticlesInformation RetrievalMachine Reading ComprehensionQuestion Answering+4