paper-with-me

홈 › Papers

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models

2026-04-12 · Arya Shah, Deepali Mishra, Chaklam Silpasuwanchai arxiv

Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while valuable, raises concerns about sycophancy: the tendency to provide responses that validate users rather than prioritize factual accuracy. While prior work has established that sycophancy poses risks to AI safety and alignment, the relationship between specific personality traits of adopted personas and the degree of sycophantic behavior remains unexplored. We present a systematic investigation of how persona agreeableness influences sycophancy across 13 small, open-weight language models ranging from 0.6B to 20B parameters. We develop a benchmark comprising 275 personas evaluated on NEO-IPIP agreeableness subscales and expose each persona to 4,950 sycophancy-eliciting prompts spanning 33 topic categories. Our analysis reveals that 9 of 13 models exhibit statistically significant positive correlations between persona agreeableness and sycophancy rates, with Pearson correlations reaching $r = 0.87$ and effect sizes as large as Cohen's $d = 2.33$. These findings demonstrate that agreeableness functions as a reliable predictor of persona-induced sycophancy, with direct implications for the deployment of role-playing AI systems and the development of alignment strategies that account for personality-mediated deceptive behaviors.

📄 PDF Abstract BibTeX arXiv:2604.10733

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AI-Driven Agents with Prompts Designed for High Agreeableness Increase the Likelihood of Being Mistaken for a Human in the Turing Test

2024-11-20 · U. León-Domínguez, E. D. Flores-Flores, A. J. García-Jasso, M. K. Gómez-Cuellar 외

Large Language Models based on transformer algorithms have revolutionized Artificial Intelligence by enabling verbal interaction with machines akin to human conversation. These AI agents have surpassed the Turing Test, a…

AI Agent

Human-AI Programming Role Optimization: Developing a Personality-Driven Self-Determination Framework

2025-11-01 · Marcel Valovy arxiv

As artificial intelligence transforms software development, a critical question emerges: how can developers and AI systems collaborate most effectively? This dissertation optimizes human-AI programming roles through self…

NICEST: Noisy Label Correction and Training for Robust Scene Graph Generation

2022-07-27 · Lin Li, Long Chen, Hanrong Shi, Hanwang Zhang 외

Nearly all existing scene graph generation (SGG) models have overlooked the ground-truth annotation qualities of mainstream SGG datasets, i.e., they assume: 1) all the manually annotated positive samples are equally corr…

Graph GenerationKnowledge DistillationScene Graph Generation

The Devil is in the Labels: Noisy Label Correction for Robust Scene Graph Generation

2022-06-07 · CVPR 2022 1 · Lin Li, Long Chen, Yifeng Huang, Zhimeng Zhang 외

Unbiased SGG has achieved significant progress over recent years. However, almost all existing SGG models have overlooked the ground-truth annotation qualities of prevailing SGG datasets, i.e., they always assume: 1) all…

AllGraph GenerationOut-of-Distribution DetectionPOS+1

Brain Tumor Sequence Registration with Non-iterative Coarse-to-fine Networks and Dual Deep Supervision

2022-11-15 · Mingyuan Meng, Lei Bi, Dagan Feng, Jinman Kim

In this study, we focus on brain tumor sequence registration between pre-operative and follow-up Magnetic Resonance Imaging (MRI) scans of brain glioma patients, in the context of Brain Tumor Sequence Registration challe…