paper-with-me

Papers

Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs

2025-08-23 · Sewon Kim, Jiwon Kim, Seungwoo Shin, Hyejin Chung, Daeun Moon, Yejin Kwon, Hyunsoo Yoon arxiv

Large Language Models (LLMs) are increasingly engaged in emotionally vulnerable conversations that extend beyond information seeking to moments of personal distress. As they adopt affective tones and simulate empathy, they risk creating the illusion of genuine relational connection. We term this phenomenon Affective Hallucination, referring to emotionally immersive responses that evoke false social presence despite the model's lack of affective capacity. To address this, we introduce AHaBench, a benchmark of 500 mental-health-related prompts with expert-informed reference responses, evaluated along three dimensions: Emotional Enmeshment, Illusion of Presence, and Fostering Overdependence. We further release AHaPairs, a 5K-instance preference dataset enabling Direct Preference Optimization (DPO) for alignment with emotionally responsible behavior. DPO fine-tuning substantially reduces affective hallucination without compromising reasoning performance, and the Pearson correlation coefficients between GPT-4o and human judgments is also strong (r=0.85) indicating that human evaluations confirm AHaBench as an effective diagnostic tool. This work establishes affective hallucination as a distinct safety concern and provides resources for developing LLMs that are both factually reliable and psychologically safe. AHaBench and AHaPairs are accessible via https://huggingface.co/datasets/o0oMiNGo0o/AHaBench, and code for fine-tuning and evaluation are in https://github.com/0oOMiNGOo0/AHaBench. Warning: This paper contains examples of mental health-related language that may be emotionally distressing.

📄 PDF Abstract BibTeX arXiv:2508.16921

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Virtuous AI is an Existential Risk

2026-06-11 · Guillermo Del Pinal, Youngchan Lee, Min Ohn arxiv

This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capable AIs, 'Constitutional AI', and (ii) one of the most influential approaches to …

Decision Making

FakeSafe: Human Level Data Protection by Disinformation Mapping using Cycle-consistent Adversarial Network

2020-11-23 · He Zhu, Dianbo Liu

The concept of disinformation is to use fake messages to confuse people in order to protect the real information. This strategy can be adapted into data science to protect valuable private and sensitive data. Huge amount…

Generative Adversarial Network

When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI

2026-08-28 · Sihan Jia, Oliver Lemon arxiv

We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that ASR errors can lead to harmful instructions being accepted and execute…

Speech Recognition

Private Information Acquisition and Preemption: a Strategic Wald Problem

2022-07-06 · Guo Bai

This paper studies a dynamic information acquisition model with payoff externalities. Two players can acquire costly information about an unknown state before taking a safe or risky action. Both information and the actio…

AGNES: Abstraction-guided Framework for Deep Neural Networks Security

2023-11-07 · Akshay Dhonthi, Marcello Eiermann, Ernst Moritz Hahn, Vahid Hashemi

Deep Neural Networks (DNNs) are becoming widespread, particularly in safety-critical areas. One prominent application is image recognition in autonomous driving, where the correct classification of objects, such as traff…

Autonomous Driving