paper-with-me

홈 › Papers

Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task Performance

2025-08-27 · Pedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, Benjamin Roth arxiv

Expert persona prompting -- assigning roles such as expert in math to language models -- is widely used for task improvement. However, prior work shows mixed results on its effectiveness, and does not consider when and why personas should improve performance. We analyze the literature on persona prompting for task improvement and distill three desiderata: 1) performance advantage of expert personas, 2) robustness to irrelevant persona attributes, and 3) fidelity to persona attributes. We then evaluate 9 state-of-the-art LLMs across 27 tasks with respect to these desiderata. We find that expert personas usually lead to positive or non-significant performance changes. Surprisingly, models are highly sensitive to irrelevant persona details, with performance drops of almost 30 percentage points. In terms of fidelity, we find that while higher education, specialization, and domain-relatedness can boost performance, their effects are often inconsistent or negligible across tasks. We propose mitigation strategies to improve robustness -- but find they only work for the largest, most capable models. Our findings underscore the need for more careful persona design and for evaluation schemes that reflect the intended effects of persona usage.

📄 PDF Abstract BibTeX arXiv:2508.19764

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models

2025-09-09 · Flor Miriam Plaza-del-Arco, Paul Röttger, Nino Scherrer, Emanuele Borgonovo 외 arxiv

Large language models (LLMs) are increasingly integrated into our daily lives and personalized. However, LLM personalization might also increase unintended side effects. Recent work suggests that persona prompting can le…

Natural Language Inference

Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs

2023-11-08 · Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan 외

Recent works have showcased the ability of LLMs to embody diverse personas in their responses, exemplified by prompts like 'You are Yoda. Explain the Theory of Relativity.' While this ability allows personalization of LL…

FairnessMath

ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems

2026-02-09 · Jinnuo Liu, Chuke Liu, Hua Shen arxiv

Multi-agent large language model (LLM) systems increasingly consist of agents that observe and respond to one another's outputs. While value alignment is typically evaluated for isolated models, how value perturbations p…

Predicting Personas Using Mechanic Frequencies and Game State Traces

2022-03-24 · Michael Cerny Green, Ahmed Khalifa, M Charity, Debosmita Bhaumik 외

We investigate how to efficiently predict play personas based on playtraces. Play personas can be computed by calculating the action agreement ratio between a player and a generative model of playing behavior, a so-calle…

Persona Cartography: Charting Language Model Personality Traits in Weight Space

2026-07-08 · Luke Baines, Anton Gonzalvez Hawthorne, Mariia Koroliuk, Irakli Shalibashvili 외 arxiv

Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them. Our central insight is to tre…