paper-with-me

홈 › Papers

The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs

2025-09-03 · Pengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez arxiv

Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems, with advanced LLMs displaying consistent behavioral tendencies resembling human traits like agreeableness and self-regulation. Understanding these patterns is crucial, yet prior work primarily relied on simplified self-reports and heuristic prompting, with little behavioral validation. In this study, we systematically characterize LLM personality across three dimensions: (1) the dynamic emergence and evolution of trait profiles throughout training stages; (2) the predictive validity of self-reported traits in behavioral tasks; and (3) the impact of targeted interventions, such as persona injection, on both self-reports and behavior. Our findings reveal that instructional alignment (e.g., RLHF, instruction tuning) significantly stabilizes trait expression and strengthens trait correlations in ways that mirror human data. However, these self-reported traits do not reliably predict behavior, and observed associations often diverge from human patterns. While persona injection successfully steers self-reports in the intended direction, it exerts little or inconsistent effect on actual behavior. By distinguishing surface-level trait expression from behavioral consistency, our findings challenge assumptions about LLM personality and underscore the need for deeper evaluation in alignment and interpretability.

📄 PDF Abstract BibTeX arXiv:2509.03730

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Seeing the Evidence, Missing the Answer: Tool-Guided Vision-Language Models on Visual Illusions

2026-03-31 · Xuesong Wang, Harry Wang arxiv

Vision-language models (VLMs) exhibit a systematic bias when confronted with classic optical illusions: they overwhelmingly predict the illusion as "real" regardless of whether the image has been counterfactually modifie…

Image ManipulationSpatial ReasoningImage Compression

Brain-inspired bodily self-perception model for robot rubber hand illusion

2023-03-22 · Yuxuan Zhao, Enmeng Lu, Yi Zeng

At the core of bodily self-consciousness is the perception of the ownership of one's body. Recent efforts to gain a deeper understanding of the mechanisms behind the brain's encoding of the self-body have led to various …

Persona-E$^2$: A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual Events

2026-04-10 · Yuqin Yang, Haowu Zhou, Haoran Tu, Zhiwen Hui 외 arxiv

Most affective computing research treats emotion as a static property of text, focusing on the writer's sentiment while overlooking the reader's perspective. This approach ignores how individual personalities lead to div…

Assessing the nature of large language models: A caution against anthropocentrism

2023-09-14 · Ann Speed

Generative AI models garnered a large amount of public attention and speculation with the release of OpenAIs chatbot, ChatGPT. At least two opinion camps exist: one excited about possibilities these models offer for fund…

Chatbot

Quadripolar Relational Model: a framework for the description of borderline and narcissistic personality disorders

2015-12-18 · Alessandro Fontana

Borderline personality disorder and narcissistic personality disorder are important nosographic entities and have been subject of intensive investigations. The currently prevailing psychodynamic theory for mental disorde…