Speaking rate attention-based duration prediction for speed control TTS
With the advent of high-quality speech synthesis, there is a lot of interest in controlling various prosodic attributes of speech. Speaking rate is an essential attribute towards modelling the expressivity of speech. In this work, we propose a novel approach to control the speaking rate for non-autoregressive TTS. We achieve this by conditioning the speaking rate inside the duration predictor, allowing implicit speaking rate control. We show the benefits of this approach by synthesising audio at various speaking rate factors and measuring the quality of speaking rate-controlled synthesised speech. Further, we study the effect of the speaking rate distribution of the training data towards effective rate control. Finally, we fine-tune a baseline pretrained TTS model to obtain speaking rate control TTS. We provide various analyses to showcase the benefits of using this proposed approach, along with objective as well as subjective metrics. We find that the proposed methods have higher subjective scores and lower speaker rate errors across many speaking rate factors over the baseline.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeSpeech SynthesisSimilar Papers 제목 키워드 기반
Attention and Encoder-Decoder based models for transforming articulatory movements at different speaking rates
While speaking at different rates, articulators (like tongue, lips) tend to move differently and the enunciations are also of different durations. In the past, affine transformation and DNN have been used to transform ar…
DecoderDub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, and speaking speed. Despite the strong tra…
Speech-to-Speech TranslationTranslationSpeech Timing in Typically Developing Mandarin-Speaking Children From Ages 3 To 4
This study aims to develop a better understanding of the speech timing development in Mandarin-speaking children from 3 to 4 years of age. Data were selected from two typically developing children. Four 50-min recordings…
Neural Network-Based Modeling of Phonetic Durations
A deep neural network (DNN)-based model has been developed to predict non-parametric distributions of durations of phonemes in specified phonetic contexts and used to explore which factors influence durations most. Major…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2Speaking style adaptation in Text-To-Speech synthesis using Sequence-to-sequence models with attention
Currently, there are increasing interests in text-to-speech (TTS) synthesis to use sequence-to-sequence models with attention. These models are end-to-end meaning that they learn both co-articulation and duration propert…
Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1