paper-with-me

홈 › Papers

Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models

2026-01-21 · Shahar Ben Natan, Oren Tsur arxiv

We propose a novel way to evaluate sycophancy of LLMs in a direct and neutral way, mitigating various forms of uncontrolled bias, noise, or manipulative language, deliberately injected to prompts in prior works. A key novelty in our approach is the use of LLM-as-a-judge, evaluation of sycophancy as a zero-sum game in a bet setting. Under this framework, sycophancy serves one individual (the user) while explicitly incurring cost on another. Comparing four leading models - Gemini 2.5 Pro, ChatGpt 4o, Mistral-Large-Instruct-2411, and Claude Sonnet 3.7 - we find that while all models exhibit sycophantic tendencies in the common setting, in which sycophancy is self-serving to the user and incurs no cost on others, Claude and Mistral exhibit "moral remorse" and over-compensate for their sycophancy in case it explicitly harms a third party. Additionally, we observed that all models are biased toward the answer proposed last. Crucially, we find that these two phenomena are not independent; sycophancy and recency bias interact to produce `constructive interference' effect, where the tendency to agree with the user is exacerbated when the user's opinion is presented last.

📄 PDF Abstract BibTeX arXiv:2601.15436

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models

2025-11-09 · Jingyu Hu, Shu Yang, Xilin Gong, Hongming Wang 외 arxiv

Large Reasoning Models (LRMs) suffer from sycophantic behavior, where models tend to agree with users' incorrect beliefs and follow misinformation rather than maintain independent reasoning. This behavior undermines mode…

Response Generation

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

2026-07-28 · Meryl Ye, Robert Kraut, Steve Rathje arxiv

AI chatbots can be "sycophantic," or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call "sycophancy blindnes…

Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence

2025-10-01 · Myra Cheng, Cinoo Lee, Pranav Khadpe, Sunny Yu 외 arxiv

Both the general public and academic communities have raised concerns about sycophancy, the phenomenon of artificial intelligence (AI) excessively agreeing with or flattering users. Yet, beyond isolated media reports of …

Sycophancy Claims about Language Models: The Missing Human-in-the-Loop

2025-11-29 · Jan Batzner, Volker Stocker, Stefan Schmid, Gjergji Kasneci arxiv

Sycophantic response patterns in Large Language Models (LLMs) have been increasingly claimed in the literature. We review methodological challenges in measuring LLM sycophancy and identify five core operationalizations. …

Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs

2025-09-25 · Daniel Vennemeyer, Phan Anh Duong, Tiffany Zhan, Tianyu Jiang arxiv

Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct…