paper-with-me

홈 › Papers

How RLHF Amplifies Sycophancy

2026-02-01 · Itai Shapira, Gerdus Benade, Ariel D. Procaccia arxiv

Large language models often exhibit increased sycophantic behavior after preference-based post-training, showing a stronger tendency to affirm a user's stated or implied belief even when this conflicts with factual accuracy or sound judgment. We present a formal analysis of how alignment from human feedback can increase this failure mode by identifying an explicit amplification mechanism that causally links optimization against a learned reward to bias in the human preference data used for alignment. We show that the direction of behavioral drift is determined by a covariance under the base policy between endorsing the belief signal in the prompt and the learned reward, and that the first-order effect reduces to a simple mean-gap condition. We then analyze reward learning from pairwise comparisons under random utility models like Bradley-Terry and characterize when bias in human annotators' preferences induces this reward gap. Next, we propose a training-time intervention designed to neutralize the amplification mechanism itself. Among all post-trained policies that prevent sycophantic behavior from increasing, we characterize the unique policy closest in KL divergence to the unconstrained post-trained policy, and derive the corresponding minimal reward correction as a closed-form agreement penalty. Computational experiments find that reward gaps are common and cause behavioral drift in all the configurations considered.

📄 PDF Abstract BibTeX arXiv:2602.01002

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma

2025-11-23 · Subramanyam Sahoo, Aman Chadha, Vinija Jain, Divya Chaudhary arxiv

Reinforcement Learning from Human Feedback (RLHF) is widely used for aligning large language models, yet practitioners face a persistent puzzle: improving safety often reduces fairness, scaling to diverse populations bec…

Reinforcement Learning

Linear Probe Penalties Reduce LLM Sycophancy

2024-12-01 · Henry Papadatos, Rachel Freedman

Large language models (LLMs) are often sycophantic, prioritizing agreement with their users over accurate or objective statements. This problematic behavior becomes more pronounced during reinforcement learning from huma…

Affective Context Amplifies Sycophancy in LLM Responses

2026-08-21 · Jiayi Li, Sanjana Menon, Brett Frischmann, Shomir Wilson 외 arxiv

As conversational companions, large language models (LLMs) often have access to users' emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interactions, where users sh…

Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate

2025-09-27 · Binwei Yao, Chao Shang, Wanyu Du, Jianfeng He 외 arxiv

Large language models (LLMs) often display sycophancy, a tendency toward excessive agreeability. This behavior poses significant challenges for multi-agent debating systems (MADS) that rely on productive disagreement to …

Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models

2026-06-09 · Shelly Bensal, Axel Magnuson, Aparna Balagopalan, Daniel M. Bikel arxiv

Persistent memory systems promise to make LLMs more helpful by storing user beliefs over time. We show they also make models less correct by amplifying sycophancy, wherein models prioritize agreement with users over accu…