paper-with-me

홈 › Papers

Every Answer Matters: Evaluating Commonsense with Probabilistic Measures

2024-06-06 · Qi Cheng, Michael Boratko, Pranay Kumar Yelugam, Tim O'Gorman, Nalini Singh, Andrew McCallum, Xiang Lorraine Li

Large language models have demonstrated impressive performance on commonsense tasks; however, these tasks are often posed as multiple-choice questions, allowing models to exploit systematic biases. Commonsense is also inherently probabilistic with multiple correct answers. The purpose of "boiling water" could be making tea and cooking, but it also could be killing germs. Existing tasks do not capture the probabilistic nature of common sense. To this end, we present commonsense frame completion (CFC), a new generative task that evaluates common sense via multiple open-ended generations. We also propose a method of probabilistic evaluation that strongly correlates with human judgments. Humans drastically outperform strong language model baselines on our dataset, indicating this approach is both a challenging and useful evaluation of machine common sense.

📄 PDF Abstract BibTeX arXiv:2406.04145

Code (1)

qxc101/probeval_cfc 공식 구현

Tasks

Common Sense ReasoningLanguage ModelingLanguage ModellingMultiple-choice

Similar Papers 제목 키워드 기반

Common to Whom? Regional Cultural Commonsense and LLM Bias in India

2026-01-22 · Sangmitra Madhusudan, Trush Shashank More, Steph Buongiorno, Renata Dividino 외 arxiv

Existing cultural commonsense benchmarks treat nations as monolithic, assuming uniform practices within national boundaries. But does cultural commonsense hold uniformly within a nation, or does it vary at the sub-nation…

Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning

2024-10-06 · Shramay Palta, Nishant Balepur, Peter Rankel, Sarah Wiegreffe 외

Questions involving commonsense reasoning about everyday situations often admit many $\textit{possible}$ or $\textit{plausible}$ answers. In contrast, multiple-choice question (MCQ) benchmarks for commonsense reasoning r…

Multiple-choice

SocialIQA: Commonsense Reasoning about Social Interactions

2019-04-22 · Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras 외

We introduce Social IQa, the first largescale benchmark for commonsense reasoning about social situations. Social IQa contains 38,000 multiple choice questions for probing emotional and social intelligence in a variety o…

Common Sense ReasoningCoreference ResolutionMultiple-choiceQuestion Answering+1

Social IQa: Commonsense Reasoning about Social Interactions

2019-11-01 · IJCNLP 2019 11 · Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras 외

We introduce Social IQa, the first large-scale benchmark for commonsense reasoning about social situations. Social IQa contains 38,000 multiple choice questions for probing emotional and social intelligence in a variety …

Multiple-choiceQuestion AnsweringTransfer Learning

LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning

2026-01-23 · Obed Junias, Maria Leonor Pacheco arxiv

Commonsense reasoning often involves evaluating multiple plausible interpretations rather than selecting a single atomic answer, yet most benchmarks rely on single-label evaluation, obscuring whether statements are joint…