paper-with-me

홈 › Papers

Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare

2026-01-26 · Clément Christophe, Wadood Mohammed Abdul, Prateek Munjal, Tathagata Raha, Ronnie Rajan, Praveenkumar Kanithi arxiv

As LLMs are increasingly integrated into clinical workflows, their tendency for sycophancy, prioritizing user agreement over factual accuracy, poses significant risks to patient safety. While existing evaluations often rely on subjective datasets, we introduce a robust framework grounded in medical MCQA with verifiable ground truths. We propose the Adjusted Sycophancy Score, a novel metric that isolates alignment bias by accounting for stochastic model instability, or "confusability". Through an extensive scaling analysis of the Qwen-3 and Llama-3 families, we identify a clear scaling trajectory for resilience. Furthermore, we reveal a counter-intuitive vulnerability in reasoning-optimized "Thinking" models: while they demonstrate high vanilla accuracy, their internal reasoning traces frequently rationalize incorrect user suggestions under authoritative pressure. Our results across frontier models suggest that benchmark performance is not a proxy for clinical reliability, and that simplified reasoning structures may offer superior robustness against expert-driven sycophancy.

📄 PDF Abstract BibTeX arXiv:2601.18334

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies

2024-06-06 · Aswin RRV, Nemika Tyagi, Md Nayem Uddin, Neeraj Varshney 외

This study explores the sycophantic tendencies of Large Language Models (LLMs), where these models tend to provide answers that match what users want to hear, even if they are not entirely correct. The motivation behind …

HallucinationKnowledge ProbingMisinformation

Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor

2026-04-30 · Petter Törnberg, Michelle Schimmel arxiv

Large language models (LLMs) are commonly evaluated for political bias based on their responses to fixed questionnaires, which typically place frontier models on the political left. A parallel literature shows that LLMs …

Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs

2025-09-25 · Daniel Vennemeyer, Phan Anh Duong, Tiffany Zhan, Tianyu Jiang arxiv

Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct…

Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models

2025-09-19 · Renjie Pi, Kehao Miao, Li Peihang, Runtao Liu 외 arxiv

Multimodal large language models (MLLMs) have demonstrated extraordinary capabilities in conducting conversations based on image inputs. However, we observe that MLLMs exhibit a pronounced form of visual sycophantic beha…

LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit

2026-04-21 · Manav Pandey arxiv

When a language model agrees with a user's false belief, is it failing to detect the error, or noticing and agreeing anyway? We show the latter. Across twelve open-weight models from five labs, spanning small to frontier…