paper-with-me

Papers

DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models

2025-02-27 · Weihao wu, Zhiwei Lin, Yixuan Zhou, Jingbei Li, Rui Niu, Qinghua Wu, Songjun Cao, Long Ma, Zhiyong Wu

Conversational speech synthesis (CSS) aims to synthesize both contextually appropriate and expressive speech, and considerable efforts have been made to enhance the understanding of conversational context. However, existing CSS systems are limited to deterministic prediction, overlooking the diversity of potential responses. Moreover, they rarely employ language model (LM)-based TTS backbones, limiting the naturalness and quality of synthesized speech. To address these issues, in this paper, we propose DiffCSS, an innovative CSS framework that leverages diffusion models and an LM-based TTS backbone to generate diverse, expressive, and contextually coherent speech. A diffusion-based context-aware prosody predictor is proposed to sample diverse prosody embeddings conditioned on multimodal conversational context. Then a prosody-controllable LM-based TTS backbone is developed to synthesize high-quality speech with sampled prosody embeddings. Experimental results demonstrate that the synthesized speech from DiffCSS is more diverse, contextually coherent, and expressive than existing CSS systems

📄 PDF Abstract BibTeX arXiv:2502.19924

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityLanguage ModelingLanguage ModellingSpeech Synthesis

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Evaluating expressive speech synthesis from audiobook corpora for conversational phrases

2012-05-01 · LREC 2012 5 · {\'E}va Sz{\'e}kely, Joao Paulo Cabral, Mohamed Abou-Zleikha, Peter Cahill 외

Audiobooks are a rich resource of large quantities of natural sounding, highly expressive speech. In our previous research we have shown that it is possible to detect different expressive voice styles represented in a pa…

ClusteringExpressive Speech SynthesisSpeech Synthesis

FCTalker: Fine and Coarse Grained Context Modeling for Expressive Conversational Speech Synthesis

2022-10-27 · Yifan Hu, Rui Liu, Guanglai Gao, Haizhou Li

Conversational Text-to-Speech (TTS) aims to synthesis an utterance with the right linguistic and affective prosody in a conversational context. The correlation between the current utterance and the dialogue history at th…

Speech Synthesistext-to-speechText to Speech

MIBURI: Towards Expressive Interactive Gesture Synthesis

2026-03-03 · M. Hamza Mughal, Rishabh Dabral, Vera Demberg, Christian Theobalt arxiv

Embodied Conversational Agents (ECAs) aim to emulate human face-to-face interaction through speech, gestures, and facial expressions. Current large language model (LLM)-based conversational agents lack embodiment and the…

Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis

2025-01-11 · Rui Liu, Zhenqi Jia, Feilong Bao, Haizhou Li

Conversational speech synthesis (CSS) aims to take the current dialogue (CD) history as a reference to synthesize expressive speech that aligns with the conversational style. Unlike CD, stored dialogue (SD) contains pres…

AttributeBenchmarkingRetrievalSpeech Synthesis

Generative Expressive Conversational Speech Synthesis

2024-07-31 · Rui Liu, Yifan Hu, Yi Ren, Xiang Yin 외

Conversational Speech Synthesis (CSS) aims to express a target utterance with the proper speaking style in a user-agent conversation setting. Existing CSS methods employ effective multi-modal context modeling techniques …

Speech Synthesis