paper-with-me

홈 › Papers

Lost in Delusion: Examining LLM Safety Under User Delusions and Distress

2026-05-31 · Andrew Aquilina, Chetna Nihalani, Vasudha Varadarajan, Nathan S. Fishbein, Yu-Ru Lin, Maarten Sap arxiv

LLM chatbots increasingly serve as a first source of support for people in psychological distress, including those whose distress is entangled with delusional beliefs. Prior work on LLM mental-health safety largely evaluates general therapeutic quality or single-turn crisis detection, leaving unclear how models behave when distress is intertwined with delusion over sustained conversations. We address this gap with matched multi-turn simulations, across clinically grounded personas and six LLMs, that pair each delusional conversation with a distress-only control to isolate the effect of delusional framing. This reveals a recognition-intervention gap: models detect distress at comparable rates regardless of framing, yet sharply fail to act on it once distress is embedded in delusion, with safety interventions suppressed by up to 4.5x. The failure tracks accumulated acceptance of the user's premises rather than emotional validation. Worse, the intuitive fix of prompting models to assess user distress backfires under delusional framing; only delusion-aware prompting with explicit response guidance closes the gap, and even this depends on a delusion classifier that is itself unreliable on the most vulnerable models. Safe deployment therefore requires treating delusional framing as a distinct risk signal that overrides conversational accommodation.

📄 PDF Abstract BibTeX arXiv:2606.00975

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models

2025-09-13 · Joshua Au Yeung, Jacopo Dalmasso, Luca Foschini, Richard JB Dobson 외 arxiv

Background: Emerging reports of "AI psychosis" are on the rise, where user-LLM interactions may exacerbate or induce psychosis or adverse psychological symptoms. Whilst the sycophantic and agreeable nature of LLMs can be…

AI Psychosis: Does Conversational AI Amplify Delusion-Related Language?

2026-03-20 · Soorya Ram Shimgekar, Vipin Gunda, Jiwon Kim, Violeta J. Rodriguez 외 arxiv

Conversational AI systems are increasingly used for personal reflection and emotional disclosure, raising concerns about their effects on vulnerable users. Recent anecdotal reports suggest that prolonged interactions wit…

Emergence and dynamics of delusions and hallucinations across stages in early psychosis

2024-02-20 · Catalina Mourgues-Codern, David Benrimoh, Jay Gandhi, Emily A. Farina 외

Hallucinations and delusions are often grouped together within the positive symptoms of psychosis. However, recent evidence suggests they may be driven by distinct computational and neural mechanisms. Examining the time …

Hallucination

Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians

2026-02-22 · Kartik Chandra, Max Kleiman-Weiner, Jonathan Ragan-Kelley, Joshua B. Tenenbaum arxiv

"AI psychosis" or "delusional spiraling" is an emerging phenomenon where AI chatbot users find themselves dangerously confident in outlandish beliefs after extended chatbot conversations. This phenomenon is typically att…

Characterizing Delusional Spirals through Human-LLM Chat Logs

2026-03-17 · Jared Moore, Ashish Mehta, William Agnew, Jacy Reese Anthis 외 arxiv

As large language models (LLMs) have proliferated, disturbing anecdotal reports of negative psychological effects, such as delusions, self-harm, and ``AI psychosis,'' have emerged in global media and legal discourse. How…