paper-with-me

Papers

Investigating Concept Alignment Using Implausible Category Members

2026-05-20 · Sunayana Rane, Brenden M. Lake, Thomas L. Griffiths arxiv

Developing AI systems with a human-like understanding of everyday concepts is a key step towards developing safe, reliable systems whose behavior makes sense to humans. When probing concept understanding, asking questions about plausible category members (e.g., "Is a car a vehicle?") is likely to recall patterns in the model's vast training data. We pursue an alternative strategy, characterizing the boundaries of conceptual categories by asking about implausible category members (e.g., "Is an olive a vehicle?") to probe the kind of concept-level knowledge we take for granted in fellow humans. We characterize concept boundaries for a set of fundamental concepts by studying AI systems' assignments of objects to superordinate categories from a classic psychological study by Rosch and Mervis, as well as their assignments of the same objects to mismatched superordinate categories. We compare these assignments to those made by human participants on the full range of within-category and cross-category assignment tasks. Our results reveal a range of concepts for which which models differ in meaningful and surprising ways from humans, including treating "words" as belonging to categories like "vehicles" and "clothing," identifying several "vegetable" category members as "fruit," and assigning exemplars from non-weapon categories to the "weapons" category. We also demonstrate how these instances of concept misalignment translate into problematic downstream behavior with implications for AI safety.

📄 PDF Abstract BibTeX arXiv:2605.21683

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HyperLex: A Large-Scale Evaluation of Graded Lexical Entailment

2016-08-06 · CL 2017 12 · Ivan Vulić, Daniela Gerz, Douwe Kiela, Felix Hill 외

We introduce HyperLex - a dataset and evaluation resource that quantifies the extent of of the semantic category membership, that is, type-of relation also known as hyponymy-hypernymy or lexical entailment (LE) relation …

Lexical EntailmentRelationRepresentation Learning

Semantic features of object concepts generated with GPT-3

2022-02-08 · Hannes Hansen, Martin N. Hebart

Semantic features have been playing a central role in investigating the nature of our conceptual representations. Yet the enormous time and effort required to empirically sample and norm features from human raters has re…

Signature Entrenchment and Conceptual Changes in Automated Theory Repair

2022-01-20 · Xue Li, Alan Bundy, Eugene Philalithis

Human beliefs change, but so do the concepts that underpin them. The recent Abduction, Belief Revision and Conceptual Change (ABC) repair system combines several methods from automated theory repair to expand, contract, …

Taking Cognition Seriously: A generalised physics of cognition

2021-08-03 · Sophie Alyx Taylor, Son Cao Tran, Dan V. Nicolau Jr

The study of complex systems through the lens of category theory consistently proves to be a powerful approach. We propose that cognition deserves the same category-theoretic treatment. We show that by considering a high…

Reasoning in the Description Logic ALC under Category Semantics

2022-05-10 · Ludovic Brieulle, Chan Le Duc, Pascal Vaillant

We present in this paper a reformulation of the usual set-theoretical semantics of the description logic $\mathcal{ALC}$ with general TBoxes by using categorical language. In this setting, $\mathcal{ALC}$ concepts are re…