paper-with-me

Papers

DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech

2022-07-03 · Keon Lee, Kyumin Park, Daeyoung Kim

The majority of current Text-to-Speech (TTS) datasets, which are collections of individual utterances, contain few conversational aspects. In this paper, we introduce DailyTalk, a high-quality conversational speech dataset designed for conversational TTS. We sampled, modified, and recorded 2,541 dialogues from the open-domain dialogue dataset DailyDialog inheriting its annotated attributes. On top of our dataset, we extend prior work as our baseline, where a non-autoregressive TTS is conditioned on historical information in a dialogue. From the baseline experiment with both general and our novel metrics, we show that DailyTalk can be used as a general TTS dataset, and more than that, our baseline can represent contextual information from DailyTalk. The DailyTalk dataset and baseline code are freely available for academic use with CC-BY-SA 4.0 license.

📄 PDF Abstract BibTeX arXiv:2207.01063

Code (1)

keonlee9420/DailyTalk 공식 구현 pytorch

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling

2026-08-14 · Xueqi Wang, Zhigang Wang, Runqing Zhang, Zhenqi Jia 외 arxiv

Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level ret…

Emotion Recognition in ConversationContrastive LearningSpeech Synthesis

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

2025-01-02 · Xize Cheng, Dongjie Fu, Xiaoda Yang, Minghui Fang 외

With the rapid development of large language models, researchers have created increasingly advanced spoken dialogue systems that can naturally converse with humans. However, these systems still struggle to handle the ful…

feature selectionSpoken Dialogue Systems

Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context Modeling

2023-12-19 · Rui Liu, Yifan Hu, Yi Ren, Xiang Yin 외

Conversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the pri…

Contrastive LearningSpeech Synthesis

Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis

2025-01-11 · Rui Liu, Zhenqi Jia, Feilong Bao, Haizhou Li

Conversational speech synthesis (CSS) aims to take the current dialogue (CD) history as a reference to synthesize expressive speech that aligns with the conversational style. Unlike CD, stored dialogue (SD) contains pres…

AttributeBenchmarkingRetrievalSpeech Synthesis

Towards Data Distillation for End-to-end Spoken Conversational Question Answering

2020-10-18 · Chenyu You, Nuo Chen, Fenglin Liu, Dongchao Yang 외

In spoken question answering, QA systems are designed to answer questions from contiguous text spans within the related speech transcripts. However, the most natural way that human seek or test their knowledge is via hum…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Conversational Question AnsweringQuestion Answering+2