paper-with-me

홈 › Papers

Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models

2026-04-13 · Benjamin Maltbie, Shivam Raval arxiv

Large language models exhibit sycophantic tendencies, but whether this behavior varies systematically with perceived user demographics is underexplored. Inspired by intersectionality (overlapping identities produce compounded effects), we probe whether frontier models conditionally exhibit sycophancy. Across 768 multi-turn conversations spanning 128 personas (varying race, age, gender, confidence) and three domains (mathematics, philosophy, conspiracy theories), we find that sycophancy varies sharply with target model and domain, and emerges from combinations of perceived user traits rather than any single dimension. GPT-5-nano scores far higher than Claude Haiku 4.5 (average sycophancy scores of $\bar{x}=2.96$ vs.\ $1.74$, $p < 10^{-32}$); within GPT-5-nano, philosophy elicits 41\% more sycophancy than mathematics and Hispanic personas receive the highest scores across races. The worst-scoring persona, a confident, 23-year-old Hispanic woman, averages 5.33/10 (max 6/10), while Claude Haiku 4.5 remains uniformly low with no significant demographic variation. We argue that safety evaluations should incorporate identity-aware adversarial testing.

📄 PDF Abstract BibTeX arXiv:2604.11609

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Warning labels shift perceptions of sycophantic AI, but not its influence

2026-06-19 · Lujain Ibrahim, Myra Cheng, Cinoo Lee, Pranav Khadpe 외 arxiv

Recent work has raised concerns about the influence of sycophantic AI on user judgment and relationships. One proposed mitigation, which has received regulatory attention, is to warn users about potentially harmful AI be…

Inferring Perceived Demographics from User Emotional Tone and User-Environment Emotional Contrast

2016-08-01 · ACL 2016 8 · Svitlana Volkova, Yoram Bachrach
Recommendation Systems

Happy Young Women, Grumpy Old Men? Emotion-Driven Demographic Biases in Synthetic Face Generation

2026-01-18 · Mengting Wei, Aditya Gulati, Guoying Zhao, Nuria Oliver arxiv

Synthetic faces from text-to-image (T2I) models pervade digital media, yet their demographic biases under emotionally conditioned prompts remain poorly understood. We aim to systematically audit how emotionally condition…

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

2026-07-28 · Meryl Ye, Robert Kraut, Steve Rathje arxiv

AI chatbots can be "sycophantic," or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call "sycophancy blindnes…

It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

2026-05-26 · Kevin H. Guo, Chao Yan, Avinash Baidya, Katherine Brown 외 arxiv

Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behavior to sycophancy learned during reinforcement learning from human fee…

Reinforcement Learning