paper-with-me

Papers

Empirically evaluating commonsense intelligence in large language models with large-scale human judgments

2025-05-15 · Tuan Dung Nguyen, Duncan J. Watts, Mark E. Whiting

Commonsense intelligence in machines is often assessed by static benchmarks that compare a model's output against human-prescribed correct labels. An important, albeit implicit, assumption of these labels is that they accurately capture what any human would think, effectively treating human common sense as homogeneous. However, recent empirical work has shown that humans vary enormously in what they consider commonsensical; thus what appears self-evident to one benchmark designer may not be so to another. Here, we propose a novel method for evaluating common sense in artificial intelligence (AI), specifically in large language models (LLMs), that incorporates empirically observed heterogeneity among humans by measuring the correspondence between a model's judgment and that of a human population. We first find that, when treated as independent survey respondents, most LLMs remain below the human median in their individual commonsense competence. Second, when used as simulators of a hypothetical population, LLMs correlate with real humans only modestly in the extent to which they agree on the same set of statements. In both cases, smaller, open-weight models are surprisingly more competitive than larger, proprietary frontier models. Our evaluation framework, which ties commonsense intelligence to its cultural basis, contributes to the growing call for adapting AI models to human collectivities that possess different, often incompatible, social stocks of knowledge.

📄 PDF Abstract BibTeX arXiv:2505.10309

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense Reasoning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Acquiring and Modelling Abstract Commonsense Knowledge via Conceptualization

2022-06-03 · Mutian He, Tianqing Fang, Weiqi Wang, Yangqiu Song

Conceptualization, or viewing entities and situations as instances of abstract concepts in mind and making inferences based on that, is a vital component in human intelligence for commonsense reasoning. Despite recent pr…

Knowledge Graphs

Multimodal UNcommonsense: From Odd to Ordinary and Ordinary to Odd

2026-02-02 · Yejin Son, Saejin Kim, Dongjun Min, Younjae Yu arxiv

Commonsense reasoning in multimodal contexts remains a foundational challenge in artificial intelligence. We introduce Multimodal UNcommonsense(MUN), a benchmark designed to evaluate models' ability to handle scenarios t…

CommonWhy: A Dataset for Evaluating Entity-Based Causal Commonsense Reasoning in Large Language Models

2026-05-13 · Armin Toroghi, Faeze Moradi Kalarde, Scott Sanner arxiv

To effectively interact with the real world, Large Language Models (LLMs) require entity-based commonsense reasoning, a challenging task that necessitates integrating factual knowledge about specific entities with common…

Graph Question Answering

Common Sense Beyond English: Evaluating and Improving Multilingual Language Models for Commonsense Reasoning

2021-06-13 · ACL 2021 5 · Bill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, Xiang Ren

Commonsense reasoning research has so far been limited to English. We aim to evaluate and improve popular multilingual language models (ML-LMs) to help advance commonsense reasoning (CSR) beyond English. We collect the M…

Common Sense ReasoningSentence

A Theoretically Grounded Benchmark for Evaluating Machine Commonsense

2022-03-23 · Henrique Santos, Ke Shen, Alice M. Mulvehill, Yasaman Razeghi 외

Programming machines with commonsense reasoning (CSR) abilities is a longstanding challenge in the Artificial Intelligence community. Current CSR benchmarks use multiple-choice (and in relatively fewer cases, generative)…

Generative Question AnsweringMultiple-choiceQuestion Answering