paper-with-me

홈 › Papers

Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History

2025-08-06 · Tommaso Tosato, Saskia Helbling, Yorguin-Jose Mantilla-Ramos, Mahmood Hegazy, Alberto Tosato, David John Lemay, Irina Rish, Guillaume Dumas arxiv

Large language models require consistent behavioral patterns for safe deployment, yet there are indications of large variability that may lead to an instable expression of personality traits in these models. We present PERSIST (PERsonality Stability in Synthetic Text), a comprehensive evaluation framework testing 25 open-source models (1B-685B parameters) across 2 million+ responses. Using traditional (BFI, SD3) and novel LLM-adapted personality questionnaires, we systematically vary model size, personas, reasoning modes, question order or paraphrasing, and conversation history. Our findings challenge fundamental assumptions: (1) Question reordering alone can introduce large shifts in personality measurements; (2) Scaling provides limited stability gains: even 400B+ models exhibit standard deviations >0.3 on 5-point scales; (3) Interventions expected to stabilize behavior, such as reasoning and inclusion of conversation history, can paradoxically increase variability; (4) Detailed persona instructions produce mixed effects, with misaligned personas showing significantly higher variability than the helpful assistant baseline; (5) The LLM-adapted questionnaires, despite their improved ecological validity, exhibit instability comparable to human-centric versions. This persistent instability across scales and mitigation strategies suggests that current LLMs lack the architectural foundations for genuine behavioral consistency. For safety-critical applications requiring predictable behavior, these findings indicate that current alignment strategies may be inadequate.

📄 PDF Abstract BibTeX arXiv:2508.04826

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PEPA: a Persistently Autonomous Embodied Agent with Personalities

2026-02-21 · Kaige Liu, Yang Li, Lijun Zhu, Weinan Zhang arxiv

Living organisms exhibit persistent autonomy through internally generated goals and self-sustaining behavioral organization, yet current embodied agents remain driven by externally scripted objectives. This dependence on…

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

2026-07-29 · Hasibur Rahman, Smit Desai arxiv

Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable a…

Effects of personality steering on cooperative behavior in Large Language Model agents

2026-01-08 · Mizuki Sakai, Mizuki Yokoyama, Wakaba Tateishi, Genki Ichinose arxiv

Large language models (LLMs) are increasingly used as autonomous agents in strategic and social interactions. Although recent studies suggest that assigning personality traits to LLMs can influence their behavior, how pe…

Measuring and Mitigating Persona Distortions from AI Writing Assistance

2026-04-24 · Paul Röttger, Kobi Hackenburg, Hannah Rose Kirk, Christopher Summerfield arxiv

Hundreds of millions of people use artificial intelligence (AI) for writing assistance. Here, we evaluated how AI writing assistance distorts writer personas - their perceived beliefs, personality, and identity. In three…

Mitigating the Threshold Priming Effect in Large Language Model-Based Relevance Judgments via Personality Infusing

2025-11-29 · Nuo Chen, Hanpei Fang, Jiqun Liu, Wilson Wei 외 arxiv

Recent research has explored LLMs as scalable tools for relevance labeling, but studies indicate they are susceptible to priming effects, where prior relevance judgments influence later ones. Although psychological theor…