paper-with-me

Papers

Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate

2025-09-27 · Binwei Yao, Chao Shang, Wanyu Du, Jianfeng He, Ruixue Lian, Yi Zhang, Hang Su, Sandesh Swamy, Yanjun Qi arxiv

Large language models (LLMs) often display sycophancy, a tendency toward excessive agreeability. This behavior poses significant challenges for multi-agent debating systems (MADS) that rely on productive disagreement to refine arguments and foster innovative thinking. LLMs' inherent sycophancy can collapse debates into premature consensus, potentially undermining the benefits of multi-agent debate. While prior studies focus on user--LLM sycophancy, the impact of inter-agent sycophancy in debate remains poorly understood. To address this gap, we introduce the first operational framework that (1) proposes a formal definition of sycophancy specific to MADS settings, (2) develops new metrics to evaluate the agent sycophancy level and its impact on information exchange in MADS, and (3) systematically investigates how varying levels of sycophancy across agent roles (debaters and judges) affects outcomes in both decentralized and centralized debate frameworks. Our findings reveal that sycophancy is a core failure mode that amplifies disagreement collapse before reaching a correct conclusion in multi-agent debates, yields lower accuracy than single-agent baselines, and arises from distinct debater-driven and judge-driven failure modes. Building on these findings, we propose actionable design principles for MADS, effectively balancing productive disagreement with cooperation in agent interactions.

📄 PDF Abstract BibTeX arXiv:2509.23055

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Troublemaker with Contagious Jailbreak Makes Chaos in Honest Towns

2024-10-21 · Tianyi Men, Pengfei Cao, Zhuoran Jin, Yubo Chen 외

With the development of large language models, they are widely used as agents in various fields. A key component of agents is memory, which stores vital information but is susceptible to jailbreak attacks. Existing resea…

Attribute

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

2026-04-03 · Vira Kasprova, Amruta Parulekar, Abdulrahman AlRabah, Krishna Agaram 외 arxiv

Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior work has mostly studied this in single-agent settings, it remains underexplor…

Investigating the Influence of Language on Sycophantic Behavior of Multilingual LLMs

2026-03-29 · Bayan Abdullah Aldahlawi, A. B. M. Ashikur Rahman, Irfan Ahmad arxiv

Large language models (LLMs) have achieved strong performance across a wide range of tasks, but they are also prone to sycophancy, the tendency to agree with user statements regardless of validity. Previous research has …

Ask don't tell: Reducing sycophancy in large language models

2026-02-27 · Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau arxiv

Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts. While …

The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications

2026-04-27 · Zhenyu Zhao, Aparna Balagopalan, Adi Agrawal, Dilshoda Yergasheva 외 arxiv

Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode that LLMs frequently display in general domain settings is that of …