paper-with-me

홈 › Papers

Localizing Persona Representations in LLMs

2025-05-30 · Celia Cintas, Miriam Rateike, Erik Miehling, Elizabeth Daly, Skyler Speakman

We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in the representation space of large language models (LLMs). Using a range of dimension reduction and pattern recognition methods, we first identify the model layers that show the greatest divergence in encoding these representations. We then analyze the activations within a selected layer to examine how specific personas are encoded relative to others, including their shared and distinct embedding spaces. We find that, across multiple pre-trained decoder-only LLMs, the analyzed personas show large differences in representation space only within the final third of the decoder layers. We observe overlapping activations for specific ethical perspectives -- such as moral nihilism and utilitarianism -- suggesting a degree of polysemy. In contrast, political ideologies like conservatism and liberalism appear to be represented in more distinct regions. These findings help to improve our understanding of how LLMs internally represent information and can inform future efforts in refining the modulation of specific human traits in LLM outputs. Warning: This paper includes potentially offensive sample statements.

📄 PDF Abstract BibTeX arXiv:2505.24539

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderDimensionality Reduction

Similar Papers 제목 키워드 기반

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

2026-07-29 · Ruikang Zhang, Shuo Wang, Qi Su arxiv

Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing…

Personality Vector: Modulating Personality of Large Language Models by Model Merging

2025-09-24 · Seungjong Sun, Seo Yeon Baek, Jang Hyun Kim arxiv

Driven by the demand for personalized AI systems, there is growing interest in aligning the behavior of large language models (LLMs) with human traits such as personality. Previous attempts to induce personality in LLMs …

Continuous Control

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

2026-06-08 · Yanyan Luo, Xue Han, Ruiqiao Bai, Xin Huang 외 arxiv

Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories. However, the mechanisms that enable personalization also expand the s…

Reinforcement Learning

Offline Reasoning for Efficient Recommendation: LLM-Empowered Persona-Profiled Item Indexing

2026-02-25 · Deogyong Kim, Junseong Lee, Jeongeun Lee, Changhoe Kim 외 arxiv

Recent advances in large language models (LLMs) offer new opportunities for recommender systems by capturing the nuanced semantics of user interests and item characteristics through rich semantic understanding and contex…

When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs

2026-01-16 · Zhongxiang Sun, Yi Zhan, Chenglei Shen, Weijie Yu 외 arxiv

Personalized large language models (LLMs) adapt model behavior to individual users to enhance user satisfaction, yet personalization can inadvertently distort factual reasoning. We show that when personalized LLMs face f…

Question Answering