paper-with-me

Papers

Brittle Minds, Fixable Activations: Understanding Belief Representations in Language Models

2024-06-25 · Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi, Andreas Bulling

Despite growing interest in Theory of Mind (ToM) tasks for evaluating language models (LMs), little is known about how LMs internally represent mental states of self and others. Understanding these internal mechanisms is critical - not only to move beyond surface-level performance, but also for model alignment and safety, where subtle misattributions of mental states may go undetected in generated outputs. In this work, we present the first systematic investigation of belief representations in LMs by probing models across different scales, training regimens, and prompts - using control tasks to rule out confounds. Our experiments provide evidence that both model size and fine-tuning substantially improve LMs' internal representations of others' beliefs, which are structured - not mere by-products of spurious correlations - yet brittle to prompt variations. Crucially, we show that these representations can be strengthened: targeted edits to model activations can correct wrong ToM inferences.

📄 PDF Abstract BibTeX arXiv:2406.17513

Code (0)

등록된 구현이 없습니다.

Tasks

Benchmarking

Similar Papers 제목 키워드 기반

Learning Triadic Belief Dynamics in Nonverbal Communication from Videos

2021-04-07 · CVPR 2021 1 · Lifeng Fan, Shuwen Qiu, Zilong Zheng, Tao Gao 외

Humans possess a unique social cognition capability; nonverbal communication can convey rich social information among agents. In contrast, such crucial social characteristics are mostly missing in the existing scene unde…

Scene Understanding

Modeling Others' Minds as Code

2025-09-29 · Kunal Jha, Aydan Yuenan Huang, Eric Ye, Natasha Jaques 외 arxiv

Accurate prediction of human behavior is essential for robust and safe human-AI collaboration. However, existing approaches for modeling people are often data-hungry and brittle because they either make unrealistic assum…

Action UnderstandingProgram Synthesis

FairMindSim: Alignment of Behavior, Emotion, and Belief in Humans and LLM Agents Amid Ethical Dilemmas

2024-10-14 · Yu Lei, Hao liu, Chengxing Xie, Songjia Liu 외

AI alignment is a pivotal issue concerning AI control and safety. It should consider not only value-neutral human preferences but also moral and ethical considerations. In this study, we introduced FairMindSim, which sim…

Computational Thought Experiments for a More Rigorous Philosophy and Science of the Mind

2024-05-14 · Iris Oved, Nikhil Krishnaswamy, James Pustejovsky, Joshua Hartshorne

We offer philosophical motivations for a method we call Virtual World Cognitive Science (VW CogSci), in which researchers use virtual embodied agents that are embedded in virtual worlds to explore questions in the field …

Philosophy

Mind the Gaps: Mixture-of-Minds for Human Simulation

2026-08-06 · Pranav Dahiya arxiv

Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this…