paper-with-me

홈 › Papers

AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence

2026-05-20 · Kate M. Lubrano, Faisal Sayed, Ankita Rathod, Akshansh, Craver Corbyn Thomas-Smith, Mark E. Whiting, Karina Nguyen arxiv

Emotional intelligence (EI), the ability to perceive, understand, and respond appropriately to others' emotional states, is central to human communication, and increasingly important to assess as LLMs assume conversational roles in everyday life. Existing EI benchmarks rely on synthetic prompts, single-turn cases, or third-party annotation. These approaches do not directly measure how models infer and respond to a participant's emotional state over the course of a real conversation. We introduce AttuneBench, a benchmark grounded in 200 genuine multi-turn human-model conversations in which participants conversed with anonymized LLMs and provided turn-by-turn annotations of their emotional state, the model's behavior, and their preferred responses. Across 11 evaluated models, we find that model rankings on emotion recognition, behavioral classification, preference prediction, and judged response quality are largely independent, indicating that emotionally intelligent behavior decomposes into separable capabilities. Preference alignment and response-quality judgments are substantially more model-discriminating than emotion-label accuracy. These results indicate that emotionally intelligent behavior requires predicting what kind of response a specific user wants in context, a distinction that aggregate scoring can obscure and that single-turn or synthetic formats cannot directly capture across turns. AttuneBench provides a framework for assessing each of these capabilities and for diagnosing model-specific strengths and failure modes in emotionally salient conversation.

📄 PDF Abstract BibTeX arXiv:2605.21739

Code (0)

등록된 구현이 없습니다.

Tasks

Emotional IntelligenceEmotion Recognition

Similar Papers 제목 키워드 기반

The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era

2026-01-09 · Zhixian Zhao, Shuiyuan Wang, Guojian Li, Hongfei Xue 외 arxiv

Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowing the gap between human-machine and huma…

Emotional Intelligence

H2HTalk: Evaluating Large Language Models as Emotional Companion

2025-07-04 · Boyang Wang, Yalun Wu, Hongcheng Guo, Zhoujun Li arxiv

As digital emotional support needs grow, Large Language Model companions offer promising authentic, always-available empathy, though rigorous evaluation lags behind model advancement. We present Heart-to-Heart Talk (H2HT…

Emotional Intelligence

EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models

2025-02-06 · He Hu, Yucheng Zhou, Lianzhong You, Hongbo Xu 외

With the integration of Multimodal large language models (MLLMs) into robotic systems and various AI applications, embedding emotional intelligence (EI) capabilities into these models is essential for enabling robots to …

BenchmarkingEmotional IntelligenceEmotion Recognition

SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

2026-06-24 · Liang-Yuan Wu, Zih-Ching Chen, Tongshuang Wu, Chao-Han Huck Yang 외 arxiv

As multimodal conversational systems increasingly engage in spoken interaction, their ability to navigate paralinguistic social cues has become a critical bottleneck for natural human-AI communication. However, existing …

Speech Emotion RecognitionEmotional Intelligence

Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation

2024-02-20 · Dongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon 외

Emotional Support Conversation (ESC) is a task aimed at alleviating individuals' emotional distress through daily conversation. Given its inherent complexity and non-intuitive nature, ESConv dataset incorporates support …

Emotional Intelligence