\textrm{DuReader}_{\textrm{vis}}: A Chinese Dataset for Open-domain Document Visual Question Answering
Open-domain question answering has been used in a wide range of applications, such as web search and enterprise search, which usually takes clean texts extracted from various formats of documents (e.g., web pages, PDFs, or Word documents) as the information source. However, designing different text extraction approaches is time-consuming and not scalable. In order to reduce human cost and improve the scalability of QA systems, we propose and study an \textbf{Open-domain} \textbf{Doc}ument \textbf{V}isual \textbf{Q}uestion \textbf{A}nswering (Open-domain DocVQA) task, which requires answering questions based on a collection of document images directly instead of only document texts, utilizing layouts and visual features additionally. Towards this end, we introduce the first Chinese Open-domain DocVQA dataset called \textrm{DuReader}_{\textrm{vis}}, containing about 15K question-answering pairs and 158K document images from the Baidu search engine. There are three main challenges in \textrm{DuReader}_{\textrm{vis}}: (1) long document understanding, (2) noisy texts, and (3) multi-span answer extraction. The extensive experiments demonstrate that the dataset is challenging. Additionally, we propose a simple approach that incorporates the layout and visual features, and the experimental results show the effectiveness of the proposed approach. The dataset and code will be publicly available at https://github.com/baidu/DuReader/tree/master/DuReader-vis.
Code (1)
Tasks
document understandingOpen-Domain Question AnsweringQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
sPEGG: high throughput eco-evolutionary simulations on commodity graphics processors
Integrating population genetics into community ecology theory is a major goal in ecology and evolution, but analyzing the resulting models is computationally daunting. Here we describe sPEGG ($\underline{\textrm{s}}\text…
Vocal Bursts Intensity PredictionAdding Circumscription to Decidable Fragments of First-Order Logic: A Complexity Rollercoaster
We study extensions of expressive decidable fragments of first-order logic with circumscription, in particular the two-variable fragment FO$^2$, its extension C$^2$ with counting quantifiers, and the guarded fragment GF.…
SentencePattern Synthesis via Complex-Coefficient Weight Vector Orthogonal Decomposition
This paper presents a new array response control scheme named complex-coefficient weight vector orthogonal decomposition ($ \textrm{C}^2\textrm{-WORD} $) and its application to pattern synthesis. The proposed $ \textrm{C…
Pattern Synthesis via Complex-Coefficient Weight Vector Orthogonal Decomposition--Part II: Robust Sidelobe Synthesis
In this paper, the complex-coefficient weight vector orthogonal decomposition ($ \textrm{C}^2\textrm{-WORD} $) algorithm proposed in Part I of this two paper series is extended to robust sidelobe control and synthesis wi…
Planted vertex cover problem on regular random graphs and nonmonotonic temperature-dependence in the supercooled region
We introduce a planted vertex cover problem on regular random graphs and study it by the cavity method of statistical mechanics. Different from conventional Ising models, the equilibrium ferromagnetic phase transition of…