Analysis of Wikipedia-based Corpora for Question Answering
This paper gives comprehensive analyses of corpora based on Wikipedia for several tasks in question answering. Four recent corpora are collected,WikiQA, SelQA, SQuAD, and InfoQA, and first analyzed intrinsically by contextual similarities, question types, and answer categories. These corpora are then analyzed extrinsically by three question answering tasks, answer retrieval, selection, and triggering. An indexing-based method for the creation of a silver-standard dataset for answer retrieval using the entire Wikipedia is also presented. Our analysis shows the uniqueness of these corpora and suggests a better use of them for statistical question answering learning.
Code (0)
등록된 구현이 없습니다.
Tasks
Question AnsweringRetrievalSimilar Papers 제목 키워드 기반
HopRetriever: Retrieve Hops over Wikipedia to Answer Complex Questions
Collecting supporting evidence from large corpora of text (e.g., Wikipedia) is of great challenge for open-domain Question Answering (QA). Especially, for multi-hop open-domain QA, scattered evidence pieces are required …
Document EmbeddingOpen-Domain Question AnsweringQuestion AnsweringRetrievalHow Additional Knowledge can Improve Natural Language Commonsense Question Answering?
Recently several datasets have been proposed to encourage research in Question Answering domains where commonsense knowledge is expected to play an important role. Recent language models such as ROBERTA, BERT and GPT tha…
ArticlesLanguage ModelingLanguage ModellingMultiple-choice+1Hybrid-SQuAD: Hybrid Scholarly Question Answering Dataset
Existing Scholarly Question Answering (QA) methods typically target homogeneous data sources, relying solely on either text or Knowledge Graphs (KGs). However, scholarly information often spans heterogeneous sources, nec…
Knowledge GraphsLanguage ModelingLanguage ModellingLarge Language Model+2Neural Arabic Question Answering
This paper tackles the problem of open domain factual Arabic question answering (QA) using Wikipedia as our knowledge source. This constrains the answer of any question to be a span of text in Wikipedia. Open domain QA f…
ArticlesInformation RetrievalMachine Reading ComprehensionMachine Translation+5Pre-trained Language Model for Biomedical Question Answering
The recent success of question answering systems is largely attributed to pre-trained language models. However, as language models are mostly pre-trained on general domain corpora such as Wikipedia, they often have diffi…
Language ModelingLanguage ModellingmodelQuestion Answering