paper-with-me

홈 › Papers

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

2025-10-26 · Li Zhou, Lutong Yu, You Lyu, Yihang Lin, Zefeng Zhao, Junyi Ao, Yuhao Zhang, Benyou Wang, Haizhou Li arxiv

Speech Language Models (SLMs) have made significant progress in spoken language understanding. Yet it remains unclear whether they can fully perceive non lexical vocal cues alongside spoken words, and respond with empathy that aligns with both emotional and contextual factors. Existing benchmarks typically evaluate linguistic, acoustic, reasoning, or dialogue abilities in isolation, overlooking the integration of these skills that is crucial for human-like, emotionally intelligent conversation. We present EchoMind, the first interrelated, multi-level benchmark that simulates the cognitive process of empathetic dialogue through sequential, context-linked tasks: spoken-content understanding, vocal-cue perception, integrated reasoning, and response generation. All tasks share identical and semantically neutral scripts that are free of explicit emotional or contextual cues, and controlled variations in vocal style are used to test the effect of delivery independent of the transcript. EchoMind is grounded in an empathy-oriented framework spanning 3 coarse and 12 fine-grained dimensions, encompassing 39 vocal attributes, and evaluated using both objective and subjective metrics. Testing 12 advanced SLMs reveals that even state-of-the-art models struggle with high-expressive vocal cues, limiting empathetic response quality. Analyses of prompt strength, speech source, and ideal vocal cue recognition reveal persistent weaknesses in instruction-following, resilience to natural speech variability, and effective use of vocal cues for empathy. These results underscore the need for SLMs that integrate linguistic content with diverse vocal cues to achieve truly empathetic conversational ability.

📄 PDF Abstract BibTeX arXiv:2510.22758

Code (0)

등록된 구현이 없습니다.

Tasks

Spoken Language UnderstandingResponse Generation

Similar Papers 제목 키워드 기반

Autonomous Reinforcement Learning of Multiple Interrelated Tasks

2019-06-04 · Vieri Giuliano Santucci, Gianluca Baldassarre, Emilio Cartoni

Autonomous multiple tasks learning is a fundamental capability to develop versatile artificial agents that can act in complex environments. In real-world scenarios, tasks may be interrelated (or "hierarchical") so that a…

Open-Ended Question Answeringreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Multilingual Multi-Figurative Language Detection

2023-05-31 · Huiyuan Lai, Antonio Toral, Malvina Nissim

Figures of speech help people express abstract concepts and evoke stronger emotions than literal expressions, thereby making texts more creative and engaging. Due to its pervasive and fundamental character, figurative la…

Language ModellingPrompt LearningSentence

A Novel Bi-directional Interrelated Model for Joint Intent Detection and Slot Filling

2019-06-30 · ACL 2019 7 · Haihong E, Peiqing Niu, Zhongfu Chen, Meina Song

A spoken language understanding (SLU) system includes two main tasks, slot filling (SF) and intent detection (ID). The joint model for the two tasks is becoming a tendency in SLU. But the bi-directional interrelated conn…

Intent DetectionSentenceslot-fillingSlot Filling+1

VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

2024-10-15 · Sijie Cheng, Kechen Fang, Yangyang Yu, Sicheng Zhou 외

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evalu…

Question AnsweringVideo Question AnsweringVideo UnderstandingVisual Grounding

MAC: A Benchmark for Multiple Attributes Compositional Zero-Shot Learning

2024-06-18 · Shuo Xu, Sai Wang, Xinyue Hu, Yutian Lin 외

Compositional Zero-Shot Learning (CZSL) aims to learn semantic primitives (attributes and objects) from seen compositions and recognize unseen attribute-object compositions. Existing CZSL datasets focus on single attribu…

AttributeCompositional Zero-Shot LearningZero-Shot Learning