paper-with-me

홈 › Papers

Controllable neural text-to-speech synthesis using intuitive prosodic features

2020-09-14 · Tuomo Raitio, Ramya Rasipuram, Dan Castellani

Modern neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the prosody of generated utterances often represents the average prosodic style of the database instead of having wide prosodic variation. Moreover, the generated prosody is solely defined by the input text, which does not allow for different styles for the same sentence. In this work, we train a sequence-to-sequence neural network conditioned on acoustic speech features to learn a latent prosody space with intuitive and meaningful dimensions. Experiments show that a model conditioned on sentence-wise pitch, pitch range, phone duration, energy, and spectral tilt can effectively control each prosodic dimension and generate a wide variety of speaking styles, while maintaining similar mean opinion score (4.23) to our Tacotron baseline (4.26).

📄 PDF Abstract BibTeX arXiv:2009.06775

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceSpeech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Highway Layer 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
BiGRU A Bidirectional GRU, or BiGRU, is a sequence processing model that consists of two GRUs. one taking the input in a forward…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Tanh Activation 설명 없음
Residual GRU A Residual GRU is a gated recurrent unit (GRU) that incorporates the idea of residual connections from…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis

2022-03-02 · Pengyu Cheng, ZhenHua Ling

In this paper, we propose a method of speaker adaption with intuitive prosodic features for statistical parametric speech synthesis. The intuitive prosodic features employed in this method include pitch, pitch range, spe…

Speech Synthesis

Controllable speech synthesis by learning discrete phoneme-level prosodic representations

2022-11-29 · Nikolaos Ellinas, Myrsini Christidou, Alexandra Vioni, June Sig Sung 외

In this paper, we present a novel method for phoneme-level prosody control of F0 and duration using intuitive discrete labels. We propose an unsupervised prosodic clustering process which is used to discretize phoneme-le…

ClusteringSpeech Synthesistext-to-speechText to Speech

Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP

2023-09-11 · Jinzuomu Zhong, Yang Li, Hui Huang, Korin Richmond 외

In expressive and controllable Text-to-Speech (TTS), explicit prosodic features significantly improve the naturalness and controllability of synthesised speech. However, manual prosody annotation is labor-intensive and i…

text-to-speechText to Speech

DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles

2024-12-04 · Jiaxuan Liu, Zhaoci Liu, Yajun Hu, Yingying Gao 외

Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model ba…

Prosody Predictiontext-to-speechText to Speech

Hierarchical prosody modeling and control in non-autoregressive parallel neural TTS

2021-10-06 · Tuomo Raitio, Jiangchuan Li, Shreyas Seshadri

Neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the synthetic speech often represents the average prosodic style of the database instead of having more ve…

text-to-speechText to Speech