paper-with-me

Papers

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

2026-07-23 · Baihui Wang, Bernard Koch arxiv

Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy as a one-dimensional failure mode. Models must distinguish when to incorporate others' perspectives from when to maintain a well-grounded moral judgment. We study the broader resistance-compliance process governing this distinction. Across three studies, we show that models' judgment revision is structured along three dimensions that parallel classic phenomena in human social psychology: the distance between an incoming view and the model's initial position, the source attribution of that view, and the coalition structure supporting it. Models are generally more receptive to nearby positions, more influenced by views presented as their own prior judgments, and differently responsive to group pressure. These findings recast sycophancy as one expression of a broader judgment-updating process shaped by social influence. Our framework provides a principled basis for distinguishing constructive belief revision from sycophantic compliance, thereby supporting better alignment in morally consequential interactions.

📄 PDF Abstract BibTeX arXiv:2607.21558

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition

2026-04-07 · Muhammad Ahmed Mohsin, Ahsan Bilal, Muhammad Umer, Emily Fox arxiv

Large language models exhibit sycophancy, the tendency to shift their stated positions toward perceived user preferences or authority cues regardless of evidence. Standard alignment methods fail to correct this because s…

Moral Sycophancy in Vision Language Models

2026-02-09 · Shadman Rabby, Md. Hefzul Hossain Papon, Sabbir Ahmed, Nokimul Hasan Arif 외 arxiv

Sycophancy in Vision-Language Models (VLMs) refers to their tendency to align with user opinions, often at the expense of moral or factual accuracy. While prior studies have explored sycophantic behavior in general conte…

What Makes AI Applications Acceptable or Unacceptable? A Predictive Moral Framework

2025-08-26 · Kimmo Eriksson, Simon Karlsson, Irina Vartanova, Pontus Strimling arxiv

As artificial intelligence rapidly transforms society, developers and policymakers struggle to anticipate which applications will face public moral resistance. We propose that these judgments are not idiosyncratic but sy…

Beyond Compliance: A Resistance-Informed Motivation Reasoning Framework for Challenging Psychological Client Simulation

2026-04-12 · Danni Liu, Bo Liu, Yuxin Hu, Hantao Zhao 외 arxiv

Psychological client simulators have emerged as a scalable solution for training and evaluating counselor trainees and psychological LLMs. Yet existing simulators exhibit unrealistic over-compliance, leaving counselors u…

Reinforcement LearningResponse Generation

Do Consumers Accept AIs as Moral Compliance Agents?

2026-03-23 · Greg Nyilasy, Abraham Ryan Ade Putra Hito, Jennifer Overbeck, Brock Bastian 외 arxiv

Consumers are generally resistant to Artificial Intelligence (AI) involvement in moral decision-making, perceiving moral agency as requiring uniquely human traits. This research investigates whether consumers might inste…