paper-with-me

홈 › Papers

Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data

2025-01-19 · Jingran Xie, Shun Lei, Yue Yu, Yang Xiang, Hui Wang, Xixin Wu, Zhiyong Wu

Empathetic dialogue is crucial for natural human-computer interaction, allowing the dialogue system to respond in a more personalized and emotionally aware manner, improving user satisfaction and engagement. The emergence of large language models (LLMs) has revolutionized dialogue generation by harnessing their powerful capabilities and shown its potential in multimodal domains. Many studies have integrated speech with text-based LLMs to take speech question as input and output text response. However, the lack of spoken question-answering datasets that include speech style information to supervised fine-tuning (SFT) limits the performance of these systems. As a result, while these systems excel at understanding speech content, they often struggle to generate empathetic responses. In response, we propose a novel approach that circumvents the need for question-answering data, called Listen, Perceive, and Express (LPE). Our method employs a two-stage training process, initially guiding the LLM to listen the content and perceive the emotional aspects of speech. Subsequently, we utilize Chain-of-Thought (CoT) prompting to unlock the model's potential for expressing empathetic responses based on listened spoken content and perceived emotional cues. We employ experiments to prove the effectiveness of proposed method. To our knowledge, this is the first attempt to leverage CoT for speech-based dialogue.

📄 PDF Abstract BibTeX arXiv:2501.10937

Code (0)

등록된 구현이 없습니다.

Tasks

Dialogue GenerationQuestion Answering

Methods 이 논문이 사용한 방법론

AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Cause-Aware Empathetic Response Generation via Chain-of-Thought Fine-Tuning

2024-08-21 · Xinhao Chen, Chong Yang, Man Lan, Li Cai 외

Empathetic response generation endows agents with the capability to comprehend dialogue contexts and react to expressed emotions. Previous works predominantly focus on leveraging the speaker's emotional labels, but ignor…

DiversityEmpathetic Response GenerationResponse Generation

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

2025-05-31 · Siddhant Arora, Jinchuan Tian, Hayato Futami, Jee-weon Jung 외

Unlike traditional cascaded pipelines, end-to-end (E2E) spoken dialogue systems preserve full differentiability and capture non-phonemic information, making them well-suited for modeling spoken interactions. However, exi…

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition+5

Reflecting Twice before Speaking with Empathy: Self-Reflective Alternating Inference for Empathy-Aware End-to-End Spoken Dialogue

2026-01-26 · Yuhang Jia, Pei Liu, Haoqin Sun, Jiaming Zhou 외 arxiv

End-to-end Spoken Language Models (SLMs) hold great potential for paralinguistic perception, and numerous studies have aimed to enhance their capabilities, particularly for empathetic dialogue. However, current approache…

Reinforcement LearningResponse Generation

FreezeEmpath: Efficient Training for Empathetic Spoken Chatbots with Frozen LLMs

2026-04-20 · Yun Hong, Yan Zhou, Yang Feng arxiv

Empathy is essential for fostering natural interactions in spoken dialogue systems, as it enables machines to recognize the emotional tone of human speech and deliver empathetic responses. Recent research has made signif…

Speech Emotion Recognition

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue

2026-06-11 · Wen Zhang, Xiaocui Yang, Zhuoyue Gao, Shi Feng 외 arxiv

Empathetic spoken dialogue systems require not only semantically appropriate responses but also emotionally aligned prosodic expression. However, cascade pipelines often discard acoustic cues during speech-to-text conver…

Dialogue GenerationResponse GenerationSpeech Synthesis