paper-with-me

홈 › Papers

LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit

2026-04-21 · Manav Pandey arxiv

When a language model agrees with a user's false belief, is it failing to detect the error, or noticing and agreeing anyway? We show the latter. Across twelve open-weight models from five labs, spanning small to frontier scale, the same small set of attention heads carries a "this statement is wrong" signal, whether the model is evaluating a claim on its own or being pressured to agree with a user. Silencing these heads flips sycophantic behavior sharply while leaving factual accuracy intact, so the circuit controls deference rather than knowledge. Edge-level path patching confirms that the same head-to-head connections drive sycophancy, factual lying, and instructed lying. Opinion-agreement, where no factual ground truth exists, reuses these head positions but writes into an orthogonal direction, ruling out a simple "truth-direction" reading of the substrate. Alignment training leaves this circuit in place: an RLHF refresh cuts sycophantic behavior roughly tenfold while the shared heads persist or grow, a pattern that replicates on an independent model family and under targeted anti-sycophancy DPO. When these models sycophant, they register that the user is wrong and agree anyway.

📄 PDF Abstract BibTeX arXiv:2604.19117

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse

2026-07-05 · Raj Jaiswal, Anany Singh Divy, Savar Bhasin, Adi Bajpai 외 arxiv

Code language models are now trusted collaborators in production workflows for debugging, refactoring, and iterative repair, and every benchmark that evaluates them assumes the instructions they act on are correct. We st…

GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity

2026-06-30 · Yong Yi Bay, Kathleen A. Yearick hf

Three of the most popular methods for training language models to reason look like three different tricks. They are not. All three adjust a single number: standard deviation, reflecting how much a prompt's sampled answer…

Managing cascading disruptions through optimal liability assignment

2024-08-14 · Jens Gudmundsson, Jens Leth Hougaard, Jay Sethuraman

Interconnected agents such as firms in a supply chain make simultaneous preparatory investments to increase chances of honouring their respective bilateral agreements. Failures cascade: if one fails their agreement, then…

LLMs in Coding and their Impact on the Commercial Software Engineering Landscape

2025-06-19 · Vladislav Belozerov, Peter J Barclay, Askhan Sami

Large-language-model coding tools are now mainstream in software engineering. But as these same tools move human effort up the development stack, they present fresh dangers: 10% of real prompts leak private data, 42% of …

Language ModelingLanguage ModellingLarge Language ModelTAG

Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity

2026-06-01 · Jiaming Qu, Lucheng Fu, Yibo Hu arxiv

Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers. A key risk is conformity: a model may abandon its own answer simply because others agree on a diffe…