paper-with-me

홈 › Papers

Hierarchical prosody modeling and control in non-autoregressive parallel neural TTS

2021-10-06 · Tuomo Raitio, Jiangchuan Li, Shreyas Seshadri

Neural text-to-speech (TTS) synthesis can generate speech that is indistinguishable from natural speech. However, the synthetic speech often represents the average prosodic style of the database instead of having more versatile prosodic variation. Moreover, many models lack the ability to control the output prosody, which does not allow for different styles for the same text input. In this work, we train a non-autoregressive parallel neural TTS front-end model hierarchically conditioned on both coarse and fine-grained acoustic speech features to learn a latent prosody space with intuitive and meaningful dimensions. Experiments show that a non-autoregressive TTS model hierarchically conditioned on utterance-wise pitch, pitch range, duration, energy, and spectral tilt can effectively control each prosodic dimension, generate a wide variety of speaking styles, and provide word-wise emphasis control, while maintaining equal or better quality to the baseline model.

📄 PDF Abstract BibTeX arXiv:2110.02952

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

Hierarchical Prosody Modeling for Non-Autoregressive Speech Synthesis

2020-11-12 · Chung-Ming Chien, Hung-Yi Lee

Prosody modeling is an essential component in modern text-to-speech (TTS) frameworks. By explicitly providing prosody features to the TTS model, the style of synthesized utterances can thus be controlled. However, predic…

Speech Synthesistext-to-speechText to Speech

Improved Prosodic Clustering for Multispeaker and Speaker-independent Phoneme-level Prosody Control

2021-11-19 · Myrsini Christidou, Alexandra Vioni, Nikolaos Ellinas, Georgios Vamvoukakis 외

This paper presents a method for phoneme-level prosody control of F0 and duration on a multispeaker text-to-speech setup, which is based on prosodic clustering. An autoregressive attention-based model is used, incorporat…

ClusteringData Augmentationtext-to-speechText to Speech

Hierarchical and Multi-Scale Variational Autoencoder for Diverse and Natural Non-Autoregressive Text-to-Speech

2022-04-08 · Jae-Sung Bae, Jinhyeok Yang, Tae-Jun Bak, Young-Sun Joo

This paper proposes a hierarchical and multi-scale variational autoencoder-based non-autoregressive text-to-speech model (HiMuV-TTS) to generate natural speech with diverse speaking styles. Recent advances in non-autoreg…

Diversitytext-to-speechText to Speech

DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles

2024-12-04 · Jiaxuan Liu, Zhaoci Liu, Yajun Hu, Yingying Gao 외

Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model ba…

Prosody Predictiontext-to-speechText to Speech

Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis

2023-03-14 · Chunyu Qiang, Peng Yang, Hao Che, Ying Zhang 외

Cross-speaker style transfer in speech synthesis aims at transferring a style from source speaker to synthesized speech of a target speaker's timbre. In most previous methods, the synthesized fine-grained prosody feature…

Prosody PredictionSpeech SynthesisStyle Transfer