paper-with-me

홈 › Papers

Accounting for Sycophancy in Language Model Uncertainty Estimation

2024-10-17 · Anthony Sicilia, Mert Inan, Malihe Alikhani

Effective human-machine collaboration requires machine learning models to externalize uncertainty, so users can reflect and intervene when necessary. For language models, these representations of uncertainty may be impacted by sycophancy bias: proclivity to agree with users, even if they are wrong. For instance, models may be over-confident in (incorrect) problem solutions suggested by a user. We study the relationship between sycophancy and uncertainty estimation for the first time. We propose a generalization of the definition of sycophancy bias to measure downstream impacts on uncertainty estimation, and also propose a new algorithm (SyRoUP) to account for sycophancy in the uncertainty estimation process. Unlike previous works on sycophancy, we study a broad array of user behaviors, varying both correctness and confidence of user suggestions to see how model answers (and their certainty) change. Our experiments across conversation forecasting and question-answering tasks show that user confidence plays a critical role in modulating the effects of sycophancy, and that SyRoUP can better predict these effects. From these results, we argue that externalizing both model and user uncertainty can help to mitigate the impacts of sycophancy bias.

📄 PDF Abstract BibTeX arXiv:2410.14746

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingmodelQuestion Answering

Similar Papers 제목 키워드 기반

It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

2026-05-26 · Kevin H. Guo, Chao Yan, Avinash Baidya, Katherine Brown 외 arxiv

Large language models (LLMs) are known to abandon their initial stance to conform to user pushback. While prior research largely attributes this behavior to sycophancy learned during reinforcement learning from human fee…

Reinforcement Learning

Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories

2025-09-20 · Mohammad Beigi, Ying Shen, Parshin Shojaee, Qifan Wang 외 arxiv

Despite the remarkable capabilities of large language models, current training paradigms inadvertently foster \textit{sycophancy}, i.e., the tendency of a model to agree with or reinforce user-provided information even w…

Reinforcement Learning

Challenging the Evaluator: LLM Sycophancy Under User Rebuttal

2025-09-20 · Sungwon Kim, Daniel Khashabi arxiv

Large Language Models (LLMs) often exhibit sycophancy, distorting responses to align with user beliefs, notably by readily agreeing with user counterarguments. Paradoxically, LLMs are increasingly adopted as successful e…

BASIL: Bayesian Assessment of Sycophancy in LLMs

2025-08-23 · Katherine Atwell, Pedram Heydari, Anthony Sicilia, Malihe Alikhani arxiv

Sycophancy (overly agreeable or flattering behavior) poses a fundamental challenge for human-AI collaboration, particularly in high-stakes decision-making domains such as health, law, and education. A central difficulty …

Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare

2026-01-26 · Clément Christophe, Wadood Mohammed Abdul, Prateek Munjal, Tathagata Raha 외 arxiv

As LLMs are increasingly integrated into clinical workflows, their tendency for sycophancy, prioritizing user agreement over factual accuracy, poses significant risks to patient safety. While existing evaluations often r…