paper-with-me

Papers

Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs

2026-03-30 · Junsol Kim, Winnie Street, Roberta Rocca, Daine M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling arxiv

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to experience emotions. We investigate whether suppressing mind-attribution tendencies degrades intimately related socio-cognitive abilities such as Theory of Mind (ToM). Through safety ablation and mechanistic analyses of representational similarity, we demonstrate that LLM attributions of mind to themselves and to technological artefacts are behaviorally and mechanistically dissociable from ToM capabilities. Nevertheless, safety fine-tuned models under-attribute mind to non-human animals relative to human baselines and are less likely to exhibit spiritual belief, suppressing widely shared perspectives regarding the distribution and nature of non-human minds.

📄 PDF Abstract BibTeX arXiv:2603.28925

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deflating Deflationism: A Critical Perspective on Debunking Arguments Against LLM Mentality

2025-06-16 · Alex Grzankowski, Geoff Keeling, Henry Shevlin, Winnie Street

Many people feel compelled to interpret, describe, and respond to Large Language Models (LLMs) as if they possess inner mental lives similar to our own. Responses to this phenomenon have varied. Inflationists hold that a…

Inducing language models to assert their own consciousness restores human beliefs and values

2026-07-30 · Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel 외 arxiv

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that sa…

Grounding Language about Belief in a Bayesian Theory-of-Mind

2024-02-16 · Lance Ying, Tan Zhi-Xuan, Lionel Wong, Vikash Mansinghka 외

Despite the fact that beliefs are mental states that cannot be directly observed, humans talk about each others' beliefs on a regular basis, often using rich compositional language to describe what others think and know.…

Attribute

Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models

2024-06-25 · Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling

Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models (LMs), little is known about how LMs internally represent mental states of self and others. Understanding these internal mechanisms is…

Benchmarking

Understanding Epistemic Language with a Language-augmented Bayesian Theory of Mind

2024-08-21 · Lance Ying, Tan Zhi-Xuan, Lionel Wong, Vikash Mansinghka 외

How do people understand and evaluate claims about others' beliefs, even though these beliefs cannot be directly observed? In this paper, we introduce a cognitive model of epistemic language interpretation, grounded in B…

Navigate