paper-with-me

Papers

PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering

2024-04-19 · Yihao Ding, Kaixuan Ren, Jiabin Huang, Siwen Luo, Soyeon Caren Han

Document Question Answering (QA) presents a challenge in understanding visually-rich documents (VRD), particularly those dominated by lengthy textual content like research journal articles. Existing studies primarily focus on real-world documents with sparse text, while challenges persist in comprehending the hierarchical semantic relations among multiple pages to locate multimodal components. To address this gap, we propose PDF-MVQA, which is tailored for research journal articles, encompassing multiple pages and multimodal information retrieval. Unlike traditional machine reading comprehension (MRC) tasks, our approach aims to retrieve entire paragraphs containing answers or visually rich document entities like tables and figures. Our contributions include the introduction of a comprehensive PDF Document VQA dataset, allowing the examination of semantically hierarchical layout structures in text-dominant documents. We also present new VRD-QA frameworks designed to grasp textual contents and relations among document layouts simultaneously, extending page-level understanding to the entire multi-page document. Through this work, we aim to enhance the capabilities of existing vision-and-language models in handling challenges posed by text-dominant documents in VRD-QA.

📄 PDF Abstract BibTeX arXiv:2404.12720

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesInformation RetrievalMachine Reading ComprehensionQuestion AnsweringReading ComprehensionRetrievalVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MF2-MVQA: A Multi-stage Feature Fusion method for Medical Visual Question Answering

2022-11-11 · Shanshan Song, Jiangyun Li, Jing Wang, Yuanxiu Cai 외

There is a key problem in the medical visual question answering task that how to effectively realize the feature fusion of language and medical images with limited datasets. In order to better utilize multi-scale informa…

Medical Visual Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

CommVQA: Situating Visual Question Answering in Communicative Contexts

2024-02-22 · Nandita Shankar Naik, Christopher Potts, Elisa Kreiss

Current visual question answering (VQA) models tend to be trained and evaluated on image-question pairs in isolation. However, the questions people ask are dependent on their informational needs and prior knowledge about…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

MVQA-68K: A Multi-dimensional and Causally-annotated Dataset with Quality Interpretability for Video Assessment

2025-09-15 · Yanyun Pu, Kehan Li, Zeyi Huang, Zhijie Zhong 외 arxiv

With the rapid advancement of video generation models such as Sora, video quality assessment (VQA) is becoming increasingly crucial for selecting high-quality videos from large-scale datasets used in pre-training. Tradit…

Video Quality AssessmentZero-shot GeneralizationVideo Generation

AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering

2025-08-25 · Kang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan 외 arxiv

The advancement of Multimodal Large Language Models (MLLMs) has driven significant progress in Visual Question Answering (VQA), evolving from Single to Multi Image VQA (MVQA). However, the increased number of images in M…

Visual Question Answering

A Dual-Attention Learning Network with Word and Sentence Embedding for Medical Visual Question Answering

2022-10-01 · Xiaofei Huang, Hongfang Gong

Research in medical visual question answering (MVQA) can contribute to the development of computeraided diagnosis. MVQA is a task that aims to predict accurate and convincing answers based on given medical images and ass…

Medical Visual Question AnsweringQuestion AnsweringSentenceSentence Embedding+4