paper-with-me

Papers

Evaluating expressive speech synthesis from audiobook corpora for conversational phrases

2012-05-01 · LREC 2012 5 · {\'E}va Sz{\'e}kely, Joao Paulo Cabral, Mohamed Abou-Zleikha, Peter Cahill, Julie Carson-Berndsen

Audiobooks are a rich resource of large quantities of natural sounding, highly expressive speech. In our previous research we have shown that it is possible to detect different expressive voice styles represented in a particular audiobook, using unsupervised clustering to group the speech corpus of the audiobook into smaller subsets representing the detected voice styles. These subsets of corpora of different voice styles reflect the various ways a speaker uses their voice to express involvement and affect, or imitate characters. This study is an evaluation of the detection of voice styles in an audiobook in the application of expressive speech synthesis. A further aim of this study is to investigate the usability of audiobooks as a language resource for expressive speech synthesis of utterances of conversational speech. Two evaluations have been carried out to assess the effect of the genre transfer: transmitting expressive speech from read aloud literature to conversational phrases with the application of speech synthesis. The first evaluation revealed that listeners have different voice style preferences for a particular conversational phrase. The second evaluation showed that it is possible for users of speech synthesis systems to learn the characteristics of a voice style well enough to make reliable predictions about what a certain utterance will sound like when synthesised using that voice style.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringExpressive Speech SynthesisSpeech Synthesis

Similar Papers 제목 키워드 기반

Text-aware and Context-aware Expressive Audiobook Speech Synthesis

2024-06-09 · Dake Guo, Xinfa Zhu, Liumeng Xue, Yongmao Zhang 외

Recent advances in text-to-speech have significantly improved the expressiveness of synthetic speech. However, a major challenge remains in generating speech that captures the diverse styles exhibited by professional nar…

Contrastive LearningLanguage ModelingLanguage ModellingSentence+3

StyleSpeech: Self-supervised Style Enhancing with VQ-VAE-based Pre-training for Expressive Audiobook Speech Synthesis

2023-12-19 · Xueyuan Chen, Xi Wang, Shaofei Zhang, Lei He 외

The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-s…

DecoderSpeech Synthesis

SynPaFlex-Corpus: An Expressive French Audiobooks Corpus dedicated to expressive speech synthesis.

2018-05-01 · LREC 2018 5 · Aghilas Sini, Damien Lolive, Ga{\"e}lle Vidal, Marie Tahon 외
Expressive Speech SynthesisSpeech SynthesisText-To-Speech Synthesis

Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis

2022-10-01 · COLING 2022 10 · Xueyuan Chen, Shun Lei, Zhiyong Wu, Dong Xu 외

Naturalness and expressiveness are crucial for audiobook speech synthesis, but now are limited by the averaged global-scale speaking style representation. In this paper, we propose an unsupervised multi-scale context-sen…

Speech Synthesistext-to-speechText to Speech

Self-supervised Context-aware Style Representation for Expressive Speech Synthesis

2022-06-25 · Yihan Wu, Xi Wang, Shaofei Zhang, Lei He 외

Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of lab…

Contrastive LearningDeep ClusteringExpressive Speech SynthesisRepresentation Learning+1