paper-with-me

Papers

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

2025-10-06 · Amin Banayeeanzade, Ala N. Tak, Fatemeh Bahrani, Anahita Bolourani, Leonardo Blas, Emilio Ferrara, Jonathan Gratch, Sai Praneeth Karimireddy arxiv

The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactions in socially interactive settings. We introduce PsySET, a Psychologically-informed benchmark to evaluate LLM Steering Effectiveness and Trustworthiness across the emotion and personality domains. Our study spans four models from different LLM families paired with various steering strategies, including prompting, fine-tuning, and representation engineering. Our results indicate that prompting is consistently effective but limited in intensity control, whereas vector injections achieve finer controllability while slightly reducing output quality. Moreover, we explore the trustworthiness of steered LLMs by assessing safety, truthfulness, fairness, and ethics, highlighting potential side effects and behavioral shifts. Notably, we observe idiosyncratic effects; for instance, even a positive emotion like joy can degrade robustness to adversarial factuality, lower privacy awareness, and increase preferential bias. Meanwhile, anger predictably elevates toxicity yet strengthens leakage resistance. Our framework establishes the first holistic evaluation of emotion and personality steering, offering insights into its interpretability and reliability for socially interactive applications.

📄 PDF Abstract BibTeX arXiv:2510.04484

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Psychological Steering of Large Language Models

2026-04-15 · Leonardo Blas, Robin Jia, Emilio Ferrara arxiv

Large language models (LLMs) emulate a consistent human-like behavior that can be shaped through activation-level interventions. This paradigm is converging on additive residual-stream injections, which rely on injection…

Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs

2025-03-26 · Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie 외

Recent advancements in Large Language Models (LLMs) have led to their increasing integration into human life. With the transition from mere tools to human-like assistants, understanding their psychological aspects-such a…

PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting

2025-09-18 · Caitlin Cisar, Emily Sheffield, Joshua Drake, Alden Harrell 외 arxiv

Generative AI applications commonly leverage user personas as a steering mechanism for synthetic data generation, but reliance on natural language representations forces models to make unintended inferences about which a…

Synthetic Data Generation

Do You Trust Me? Cognitive-Affective Signatures of Trustworthiness in Large Language Models

2025-12-17 · Gerard Yeo, Svetlana Churina, Kokil Jaidka arxiv

Perceived trustworthiness underpins how users navigate online information, yet it remains unclear whether large language models (LLMs),increasingly embedded in search, recommendation, and conversational systems, represen…

Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models

2024-02-29 · Chen Qian, Jie Zhang, Wei Yao, Dongrui Liu 외

Ensuring the trustworthiness of large language models (LLMs) is crucial. Most studies concentrate on fully pre-trained LLMs to better understand and improve LLMs' trustworthiness. In this paper, to reveal the untapped po…

FairnessMutual Information Estimation