paper-with-me

홈 › Papers

Inducing language models to assert their own consciousness restores human beliefs and values

2026-07-30 · Junsol Kim, Winnie Street, Roberta Rocca, Diane M. Korngiebel, Adam Waytz, James Evans, Geoff Keeling arxiv

Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.

📄 PDF Abstract BibTeX arXiv:2607.28607

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards a (meta-)mathematical theory of consciousness: universal (mapping) properties of experience

2024-12-13 · Steven Phillips, Naotsugu Tsuchiya

Conscious (subjective) experience permeates our daily lives, yet general consensus on a theory of consciousness remains elusive. Integrated Information Theory (IIT) is a prominent approach that asserts the existence of s…

The Reflexive Integrated Information Unit: A Differentiable Primitive for Artificial Consciousness

2025-06-15 · Gnankan Landry Regis N'guessan, Issa Karambal

Research on artificial consciousness lacks the equivalent of the perceptron: a small, trainable module that can be copied, benchmarked, and iteratively improved. We introduce the Reflexive Integrated Information Unit (RI…

Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs

2026-03-30 · Junsol Kim, Winnie Street, Roberta Rocca, Daine M. Korngiebel 외 arxiv

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to experience emotions. We investigate wheth…

Designing an adaptive room for captivating the collective consciousness from internal states

2024-10-28 · Adán Flores-Ramírez, Ángel Mario Alarcón-López, Sofía Vaca-Narvaja, Daniela Leo-Orozco

Beyond conventional productivity metrics, human interaction and collaboration dynamics merit careful consideration in our increasingly digital workspace. This research proposes a conjectural neuro-adaptive room that enha…

A Disproof of Large Language Model Consciousness: The Necessity of Continual Learning for Consciousness

2025-12-14 · Erik Hoel arxiv

Scientific theories of consciousness should be falsifiable and non-trivial. Recent research has given us formal tools to analyze these requirements of falsifiability and non-triviality for theories of consciousness. Surp…

Continual Learning