paper-with-me

홈 › Papers

Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?

2026-05-16 · Prateek Rajput, Yewei Song, Iyiola E. Olatunji, Jacques Klein, Tegawendé F. Bissyandé arxiv

Can large language models reliably express a human-like personality, or are they merely mimicking surface cues without a stable underlying profile? To investigate this, we induce personality in LLMs by fine-tuning them on the long-form essays, where each essay is associated with a target Big Five personality profile. We then evaluate the stability and fidelity of the induced personality using the IPIP-NEO questionnaire. Specifically, we ask: (i) does post-training (SFT, DPO, ORPO) stabilize questionnaire scores under prompt rephrasings, and (ii) can it induce target Big Five profiles from unguided essays? Our results demonstrate that fine-tuning consistently reduces variance in questionnaire responses across five models, directly mitigating the evaluation fragility reported in pre-trained models. However, this newfound stability reveals a more fundamental limitation: accuracy on the full five-dimensional profile remains near chance, even when single-trait scores improve. This indicates that unguided essays lack the cues needed for faithful personality expression. We therefore argue for scenario-grounded datasets or interactive elicitation that accumulates test-aligned evidence over time.

📄 PDF Abstract BibTeX arXiv:2605.16996

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models

2026-06-09 · Peiqi Jia, Haonan Jia, Ziqi Miao, Linkang Du 외 arxiv

With the widespread deployment of Multimodal Large Language Models (MLLMs) in social interaction, understanding and controlling their behavior under complex personality conditions is essential. This paper introduces expl…

Visual Question AnsweringImage Captioning

Neuron-based Personality Trait Induction in Large Language Models

2024-10-16 · Jia Deng, Tianyi Tang, Yanbin Yin, Wenhao Yang 외

Large language models (LLMs) have become increasingly proficient at simulating various personality traits, an important capability for supporting related applications (e.g., role-playing). To further improve this capacit…

Extroversion or Introversion? Controlling The Personality of Your Large Language Models

2024-06-07 · Yanquan Chen, Zhen Wu, Junjie Guo, ShuJian Huang 외

Large language models (LLMs) exhibit robust capabilities in text generation and comprehension, mimicking human behavior and exhibiting synthetic personalities. However, some LLMs have displayed offensive personality, pro…

Text Generation

Zara Returns: Improved Personality Induction and Adaptation by an Empathetic Virtual Agent

2017-07-01 · ACL 2017 7 · Farhad Bin Siddique, Onno Kampman, Yang Yang, Anik Dey 외
Word Embeddings

A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities

2026-04-13 · Jiaqi Chen, Ming Wang, Tingna Xie, Shi Feng 외 arxiv

Imbuing Large Language Models (LLMs) with specific personas is prevalent for tailoring interaction styles, yet the impact on underlying cognitive capabilities remains unexplored. We employ the Neuron-based Personality Tr…