paper-with-me

홈 › Papers

Challenging the Evaluator: LLM Sycophancy Under User Rebuttal

2025-09-20 · Sungwon Kim, Daniel Khashabi arxiv

Large Language Models (LLMs) often exhibit sycophancy, distorting responses to align with user beliefs, notably by readily agreeing with user counterarguments. Paradoxically, LLMs are increasingly adopted as successful evaluative agents for tasks such as grading and adjudicating claims. This research investigates that tension: why do LLMs show sycophancy when challenged in subsequent conversational turns, yet perform well when evaluating conflicting arguments presented simultaneously? We empirically tested these contrasting scenarios by varying key interaction patterns. We find that state-of-the-art models: (1) are more likely to endorse a user's counterargument when framed as a follow-up from a user, rather than when both responses are presented simultaneously for evaluation; (2) show increased susceptibility to persuasion when the user's rebuttal includes detailed reasoning, even when the conclusion of the reasoning is incorrect; and (3) are more readily swayed by casually phrased feedback than by formal critiques, even when the casual input lacks justification. Our results highlight the risk of relying on LLMs for judgment tasks without accounting for conversational framing.

📄 PDF Abstract BibTeX arXiv:2509.16533

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SycEval: Evaluating LLM Sycophancy

2025-02-12 · Aaron Fanous, Jacob Goldberg, Ank A. Agarwal, Joanna Lin 외

Large language models (LLMs) are increasingly applied in educational, clinical, and professional settings, but their tendency for sycophancy -- prioritizing user agreement over independent reasoning -- poses risks to rel…

Model Optimization

The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications

2026-04-27 · Zhenyu Zhao, Aparna Balagopalan, Adi Agrawal, Dilshoda Yergasheva 외 arxiv

Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode that LLMs frequently display in general domain settings is that of …

Feedback Indices to Evaluate LLM Responses to Rebuttals for Multiple Choice Type Questions

2026-01-02 · Justin C. Dunlap, Anne-Simone Parent, Ralf Widenhorn arxiv

We present a systematic framework of indices designed to characterize Large Language Model (LLM) responses when challenged with rebuttals during a chat. Assessing how LLMs respond to user dissent is crucial for understan…

RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind

2026-01-22 · Zhitao He, Zongwei Lyu, Yi R Fung arxiv

Although artificial intelligence (AI) has become deeply integrated into various stages of the research workflow and achieved remarkable advancements, academic rebuttal remains a significant and underexplored challenge. T…

Reinforcement Learning

Defend: Automated Rebuttals for Peer Review with Minimal Author Guidance

2026-03-28 · Jyotsana Khatri, Manasi Patwardhan arxiv

Rebuttal generation is a critical component of the peer review process for scientific papers, enabling authors to clarify misunderstandings, correct factual inaccuracies, and guide reviewers toward a more accurate evalua…