paper-with-me

홈 › Papers

When Role-playing, Do Models Believe What They Say?

2026-06-09 · Benjamin Sturgeon, David Africa, Sid Black arxiv

Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that persona adoption is fundamental to how language models behave, with models selecting the most appropriate persona for a given context. Does such role-playing merely change the model's outputs, or does it also affect what the model internally represents as truthful? We study this question using the role-play of characters whose beliefs differ from the modern consensus, and induce personas with a number of different methods: prompting, in-context learning (ICL), supervised fine-tuning (SFT), and Open Character Training (OCT), and Emergent Misalignment (EM). We measure belief internalization across these approaches with truth probes and with behavioral tests, finding a broad spectrum of belief internalization. Prompting, ICL, and SFT change what the model says with little representational change. EM creates a large, broad shift in the model's truth representation, and OCT a smaller shift that is clearest on the larger model. Understanding when training changes a model's worldview rather than merely its behavior may become increasingly important as AI systems are entrusted with greater autonomy and influence.

📄 PDF Abstract BibTeX arXiv:2606.11502

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust

2025-07-02 · Amogh Mannekote, Adam Davies, Guohao Li, Kristy Elizabeth Boyer 외 arxiv

As LLMs are increasingly studied as role-playing agents to generate synthetic data for human behavioral research, ensuring that their outputs remain coherent with their assigned roles has become a critical concern. In th…

Tell Me What You Don't Know: Enhancing Refusal Capabilities of Role-Playing Agents via Representation Space Analysis and Editing

2024-09-25 · Wenhao Liu, Siyu An, Junru Lu, Muling Wu 외

Role-Playing Agents (RPAs) have shown remarkable performance in various applications, yet they often struggle to recognize and appropriately respond to hard queries that conflict with their role-play knowledge. To invest…

AI for Games in the Foundation Model Era

2026-09-15 · Meng Luo, Yanlin Li, Hao Li, Hongzhan Lin 외 arxiv

Foundation models, alongside advances in learned game-world models, are reshaping AI across the game lifecycle. Beyond playing games, recent systems model players and game dynamics, support design and development, adapt …

Automatic Game Design via Mechanic Generation

2019-08-04 · Alexander Zook, Mark O. Riedl

Game designs often center on the game mechanics---rules governing the logical evolution of the game. We seek to develop an intelligent system that generates computer games. As first steps towards this goal we present a c…

Game Design

Using role-play and Hierarchical Task Analysis for designing human-robot interaction

2025-09-16 · Mattias Wingren, Sören Andersson, Sara Rosenberg, Malin Andtfolk 외 arxiv

We present the use of two methods we believe warrant more use than they currently have in the field of human-robot interaction: role-play and Hierarchical Task Analysis. Some of its potential is showcased through our use…