ProtoQA: A Question Answering Dataset for Prototypical Common-Sense Reasoning
Given questions regarding some prototypical situation such as Name something that people usually do before they leave the house for work? a human can easily answer them via acquired experiences. There can be multiple right answers for such questions, with some more common for a situation than others. This paper introduces a new question answering dataset for training and evaluating common sense reasoning capabilities of artificial intelligence systems in such prototypical situations. The training set is gathered from an existing set of questions played in a long-running international game show FAMILY- FEUD. The hidden evaluation set is created by gathering answers for each question from 100 crowd-workers. We also propose a generative evaluation task where a model has to output a ranked list of answers, ideally covering all prototypical answers for a question. After presenting multiple competitive baseline models, we find that human performance still exceeds model scores on all evaluation metrics with a meaningful gap, supporting the challenging nature of the task.
Code (1)
Tasks
Common Sense ReasoningQuestion AnsweringSimilar Papers 제목 키워드 기반
KEPR: Knowledge Enhancement and Plausibility Ranking for Generative Commonsense Question Answering
Generative commonsense question answering (GenCQA) is a task of automatically generating a list of answers given a question. The answer list is required to cover all reasonable answers. This presents the considerable cha…
Passage RetrievalQuestion AnsweringRetrievalLarge Language Models Are Also Good Prototypical Commonsense Reasoners
Commonsense reasoning is a pivotal skill for large language models, yet it presents persistent challenges in specific tasks requiring this competence. Traditional fine-tuning approaches can be resource-intensive and pote…
StrategyQAGraDA: Graph Generative Data Augmentation for Commonsense Reasoning
Recent advances in commonsense reasoning have been fueled by the availability of large-scale human annotated datasets. Manual annotation of such datasets, many of which are based on existing knowledge bases, is expensive…
Data AugmentationHellaSwagKnowledge GraphsProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering
Visual Question Answering (VQA) is increasingly used in diverse applications ranging from general visual reasoning to safety-critical domains such as medical imaging and autonomous systems, where models must provide not …
Visual Question AnsweringVisual ReasoningDon't Take the Easy Way Out: Ensemble Based Methods for Avoiding Known Dataset Biases
State-of-the-art models often make use of superficial patterns in the data that do not generalize well to out-of-domain or adversarial settings. For example, textual entailment models often learn that particular key word…
Natural Language InferenceQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)