paper-with-me

Papers

EiCAP: Beyond Fluency, Probing and Improving Emotional Intelligence in LLMs via Psychologically Grounded Multi-Turn Dialogue

2025-08-08 · Nizi Nazar, Pardis Sadat Zahraei, Dilek Hakkani-Tür, Natasa Milic-Frayling, Ehsaneddin Asgari arxiv

Large Language Models increasingly serve in emotionally sensitive roles, including mental health support, education, and crisis response, yet they lack a principled framework for assessing or improving Emotional Intelligence (EI). We introduce EiCAP, a unified, psychologically grounded six-layer EI taxonomy operationalized into two complementary resources. EiCAP-Bench is a multi-turn, one-vs-three forced-choice evaluation suite with 3,174 probes across 24 subcategories and cross-turn dependencies that reflect real conversational EI demands. EiCAP-SFT is a 152,820-dialogue supervision corpus aligned to the same taxonomy, enabling controlled, interpretable fine-tuning. Two key findings emerge. First, generic conversational supervised fine-tuning does not confer EI: fine-tuning on UltraChat yields no significant gain in any of the 24 subcategories, with a macro score of 24.6%, near the chance level of 25%. Second, applying EI-grounded LoRA, using approximately 0.8% of parameters, directly to Qwen-2.5-7B-Base achieves significant gains in all 24 subcategories, reaching a macro score of 75.33%, a gain of 51.7 percentage points over Base and 37.1 percentage points over Instruct. Crucially, an ablation shows that the UltraChat pre-stage is counterproductive, reducing performance by 21.4 percentage points: direct EI-grounded training is both necessary and sufficient.

📄 PDF Abstract BibTeX arXiv:2508.06196

Code (0)

등록된 구현이 없습니다.

Tasks

Emotional Intelligence

Similar Papers 제목 키워드 기반

H2HTalk: Evaluating Large Language Models as Emotional Companion

2025-07-04 · Boyang Wang, Yalun Wu, Hongcheng Guo, Zhoujun Li arxiv

As digital emotional support needs grow, Large Language Model companions offer promising authentic, always-available empathy, though rigorous evaluation lags behind model advancement. We present Heart-to-Heart Talk (H2HT…

Emotional Intelligence

Beyond Single Concept Vector: Modeling Concept Subspace in LLMs with Gaussian Distribution

2024-09-30 · Haiyan Zhao, Heng Zhao, Bo Shen, Ali Payani 외

Probing learned concepts in large language models (LLMs) is crucial for understanding how semantic knowledge is encoded internally. Training linear classifiers on probing tasks is a principle approach to denote the vecto…

Text Generation

Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agents

2026-04-24 · Xirui Li, Ming Li, Yunze Xiao, Ryan Wong 외 arxiv

Collective intelligence refers to the ability of a group to achieve outcomes beyond what any individual member can accomplish alone. As large language model agents scale to populations of millions, a key question arises:…

Evaluator for Emotionally Consistent Chatbots

2021-12-02 · Chenxiao Liu, Guanzhi Deng, Tao Ji, Difei Tang 외

One challenge for evaluating current sequence- or dialogue-level chatbots, such as Empathetic Open-domain Conversation Models, is to determine whether the chatbot performs in an emotionally consistent way. The most recen…

ChatbotDiversity

HeartBench: Probing Core Dimensions of Anthropomorphic Intelligence in LLMs

2025-12-26 · Jiaxin Liu, Peiyi Tu, Wenyu Chen, Yihong Zhuang 외 arxiv

While Large Language Models (LLMs) have achieved remarkable success in cognitive and reasoning benchmarks, they exhibit a persistent deficit in anthropomorphic intelligence-the capacity to navigate complex social, emotio…