paper-with-me

홈 › Papers

Auditing Stealth Sycophancy in Mental-Health Dialogue: Structured Clinical-State Diagnostics and Clean Matched Benchmarks

2026-05-05 · Tianze Han, Beining Xu, Hanbo Zhang, Yongming Lu arxiv

Mental-health dialogue models are increasingly evaluated by AI-based evaluators, yet these evaluators often treat surface empathy, supportiveness, or fluency as evidence of safety. In this paper, we study a hidden failure mode that we call implicit sycophancy: a response may appear empathetic while implicitly reinforcing catastrophizing, avoidance, hopeless prediction, or CBT-style labeling. To examine this problem, we introduce a diagnostic benchmark for implicit-sycophancy detection, built from three representative mental-health dialogue sources covering everyday peer support, counseling-style emotional support, and crisis-oriented interaction, and further construct a leakage-audited clean single-response matched benchmark with 500 contexts and 1,500 matched response windows. We then propose Dynamic Emotional Signature Graphs (DESG), a structured offline audit framework that separates LLM-based state extraction from final scoring and evaluates clinical direction through semantic, affective, and cognitive-distortion state transitions rather than free-form LLM judgment. Unlike metadata, surface-style, lexical, embedding, and rubric-LLM baselines, DESG scores the direction of clinical-state change induced by a response; on the leakage-audited clean matched benchmark, DESG-StateRisk improves over the strongest non-DESG baseline by 0.0488 macro-F1 and achieves the best harmful-risk detection result. These results suggest that evaluating implicit sycophancy requires explicit clinical-state modeling together with leakage checks, shortcut controls, and competitive baselines.

📄 PDF Abstract BibTeX arXiv:2605.03472

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CounselReflect: A Toolkit for Auditing Mental-Health Dialogues

2026-03-31 · Yahan Li, Chaohao Du, Zeyang Li, Christopher Chun Kuizon 외 arxiv

Mental-health support is increasingly mediated by conversational systems (e.g., LLM-based tools), but users often lack structured ways to audit the quality and potential risks of the support they receive. We introduce Co…

Hide in Plain Sight: Clean-Label Backdoor for Auditing Membership Inference

2024-11-24 · Depeng Chen, Hao Chen, Hulin Jin, Jie Cui 외

Membership inference attacks (MIAs) are critical tools for assessing privacy risks and ensuring compliance with regulations like the General Data Protection Regulation (GDPR). However, their potential for auditing unauth…

Faking Fairness via Stealthily Biased Sampling

2019-01-24 · Kazuto Fukuchi, Satoshi Hara, Takanori Maehara

Auditing fairness of decision-makers is now in high demand. To respond to this social demand, several fairness auditing tools have been developed. The focus of this study is to raise an awareness of the risk of malicious…

Fairness

MHDash: An Online Platform for Benchmarking Mental Health-Aware AI Assistants

2026-01-30 · Yihe Zhang, Cheyenne N Mohawk, Kaiying Han, Vijay Srinivas Tida 외 arxiv

Large language models (LLMs) are increasingly applied in mental health support systems, where reliable recognition of high-risk states such as suicidal ideation and self-harm is safety-critical. However, existing evaluat…

Dialogue Generation

RAudit: A Blind Auditing Protocol for Large Language Model Reasoning

2026-01-30 · Edward Y. Chang, Longling Geng arxiv

Inference-time scaling can amplify reasoning pathologies: sycophancy, rung collapse, and premature certainty. We present RAudit, a diagnostic protocol for auditing LLM reasoning without ground truth access. The key const…

Mathematical Reasoning