paper-with-me

Papers

Document Visual Question Answering Challenge 2020

2020-08-20 · Minesh Mathew, Ruben Tito, Dimosthenis Karatzas, R. Manmatha, C. V. Jawahar

This paper presents results of Document Visual Question Answering Challenge organized as part of "Text and Documents in the Deep Learning Era" workshop, in CVPR 2020. The challenge introduces a new problem - Visual Question Answering on document images. The challenge comprised two tasks. The first task concerns with asking questions on a single document image. On the other hand, the second task is set as a retrieval task where the question is posed over a collection of images. For the task 1 a new dataset is introduced comprising 50,000 questions-answer(s) pairs defined over 12,767 document images. For task 2 another dataset has been created comprising 20 questions over 14,362 document images which share the same document template.

📄 PDF Abstract BibTeX arXiv:2008.08899

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringRetrievalTask 2Visual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends

2025-01-04 · Camille Barboule, Benjamin Piwowarski, Yoan Chabot

Using Large Language Models (LLMs) for Visually-rich Document Understanding (VrDU) has significantly improved performance on tasks requiring both comprehension and generation, such as question answering, albeit introduci…

document understandingQuestion AnsweringSurvey

JDocQA: Japanese Document Question Answering Dataset for Generative Language Models

2024-03-28 · Eri Onami, Shuhei Kurita, Taiki Miyanishi, Taro Watanabe

Document question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites, and it is a truly demanding task as paper and electronic forms of documents are so common i…

HallucinationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

DCQA: Document-Level Chart Question Answering towards Complex Reasoning and Common-Sense Understanding

2023-10-29 · Anran Wu, Luwei Xiao, Xingjiao Wu, Shuwen Yang 외

Visually-situated languages such as charts and plots are omnipresent in real-world documents. These graphical depictions are human-readable and are often analyzed in visually-rich documents to address a variety of questi…

Answer GenerationChart Question AnsweringCommon Sense ReasoningDocument Layout Analysis+3

AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings

2025-08-19 · Haoxuan Li, Wei Song, Aofan Liu, Peiwu Qin arxiv

Document Visual Question Answering (Document VQA) faces significant challenges when processing long documents in low-resource environments due to context limitations and insufficient training data. This paper presents Ad…

Visual Question AnsweringData AugmentationText Retrieval

\textrm{DuReader}_{\textrm{vis}}: A Chinese Dataset for Open-domain Document Visual Question Answering

2022-05-01 · Findings (ACL) 2022 5 · Le Qi, Shangwen Lv, Hongyu Li, Jing Liu 외

Open-domain question answering has been used in a wide range of applications, such as web search and enterprise search, which usually takes clean texts extracted from various formats of documents (e.g., web pages, PDFs, …

document understandingOpen-Domain Question AnsweringQuestion AnsweringVisual Question Answering+1