paper-with-me

홈 › Papers

DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset

2025-05-26 · Alkis Koudounas, Moreno La Quatra, Elena Baralis

Recent advances in conversational AI have demonstrated impressive capabilities in single-turn responses, yet multi-turn dialogues remain challenging for even the most sophisticated language models. Current dialogue datasets are limited in their emotional range, domain diversity, turn depth, and are predominantly text-only, hindering progress in developing more human-like conversational systems across modalities. To address these limitations, we present DeepDialogue, a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. Our approach pairs 9 different language models (4B-72B parameters) to generate 65,600 initial conversations, which we then evaluate through a combination of human annotation and LLM-based quality filtering. The resulting dataset reveals fundamental insights: smaller models fail to maintain coherence beyond 6 dialogue turns; concrete domains (e.g., "cars," "travel") yield more meaningful conversations than abstract ones (e.g., "philosophy"); and cross-model interactions produce more coherent dialogues than same-model conversations. A key contribution of DeepDialogue is its speech component, where we synthesize emotion-consistent voices for all 40,150 dialogues, creating the first large-scale open-source multimodal dialogue dataset that faithfully preserves emotional context across multi-turn conversations.

📄 PDF Abstract BibTeX arXiv:2505.19978

Code (0)

등록된 구현이 없습니다.

Tasks

Philosophy

Similar Papers 제목 키워드 기반

A Multi-Turn Emotionally Engaging Dialog Model

2019-08-15 · Yubo Xie, Ekaterina Svikhnushina, Pearl Pu

Open-domain dialog systems (also known as chatbots) have increasingly drawn attention in natural language processing. Some of the recent work aims at incorporating affect information into sequence-to-sequence neural dial…

ChatbotmodelOpen-Domain Dialog

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning

2026-05-30 · Sukru Samet Dindar, Riki Shimizu, Xilin Jiang, Nima Mesgarani arxiv

Empathetic spoken dialogue systems must infer a user's emotional state to respond appropriately, yet everyday speech often carries weak, neutral, or ambiguous affective cues. To address this, we introduce Sympatheia, a s…

Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech

2025-10-29 · Pedro Corrêa, João Lima, Victor Moreno, Lucas Ueda 외 arxiv

Advancements in spoken language processing have driven the development of spoken language models (SLMs), designed to achieve universal audio understanding by jointly learning text and audio representations for a wide ran…

Speech Emotion Recognition

SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

2026-06-24 · Liang-Yuan Wu, Zih-Ching Chen, Tongshuang Wu, Chao-Han Huck Yang 외 arxiv

As multimodal conversational systems increasingly engage in spoken interaction, their ability to navigate paralinguistic social cues has become a critical bottleneck for natural human-AI communication. However, existing …

Speech Emotion RecognitionEmotional Intelligence

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

2025-05-05 · Yemin Shi, Yu Shu, Siwei Dong, Guangyi Liu 외

A voice AI agent that blends seamlessly into daily life would interact with humans in an autonomous, real-time, and emotionally expressive manner. Rather than merely reacting to commands, it would continuously listen, re…

AI AgentAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Rhythm+4