paper-with-me

홈 › Papers

Speaking rate attention-based duration prediction for speed control TTS

2023-10-13 · Jesuraj Bandekar, Sathvik Udupa, Abhayjeet Singh, Anjali Jayakumar, Deekshitha G, Sandhya Badiger, Saurabh Kumar, Pooja VH, Prasanta Kumar Ghosh

With the advent of high-quality speech synthesis, there is a lot of interest in controlling various prosodic attributes of speech. Speaking rate is an essential attribute towards modelling the expressivity of speech. In this work, we propose a novel approach to control the speaking rate for non-autoregressive TTS. We achieve this by conditioning the speaking rate inside the duration predictor, allowing implicit speaking rate control. We show the benefits of this approach by synthesising audio at various speaking rate factors and measuring the quality of speaking rate-controlled synthesised speech. Further, we study the effect of the speaking rate distribution of the training data towards effective rate control. Finally, we fine-tune a baseline pretrained TTS model to obtain speaking rate control TTS. We provide various analyses to showcase the benefits of using this proposed approach, along with objective as well as subjective metrics. We find that the proposed methods have higher subjective scores and lower speaker rate errors across many speaking rate factors over the baseline.

📄 PDF Abstract BibTeX arXiv:2310.08846

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeSpeech Synthesis

Similar Papers 제목 키워드 기반

Attention and Encoder-Decoder based models for transforming articulatory movements at different speaking rates

2020-06-04 · Abhayjeet Singh, Aravind Illa, Prasanta Kumar Ghosh

While speaking at different rates, articulators (like tongue, lips) tend to move differently and the enunciations are also of different durations. In the past, affine transformation and DNN have been used to transform ar…

Decoder

Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing

2025-05-27 · Jeongsoo Choi, Jaehun Kim, Joon Son Chung

This paper introduces a cross-lingual dubbing system that translates speech from one language to another while preserving key characteristics such as duration, speaker identity, and speaking speed. Despite the strong tra…

Speech-to-Speech TranslationTranslation

Speech Timing in Typically Developing Mandarin-Speaking Children From Ages 3 To 4

2022-11-01 · ROCLING 2022 11 · Jeng Man Lew, Li-mei Chen, Yu Ching Lin

This study aims to develop a better understanding of the speech timing development in Mandarin-speaking children from 3 to 4 years of age. Data were selected from two typically developing children. Four 50-min recordings…

Neural Network-Based Modeling of Phonetic Durations

2019-09-06 · Xizi Wei, Melvyn Hunt, Adrian Skilling

A deep neural network (DNN)-based model has been developed to predict non-parametric distributions of durations of phonemes in specified phonetic contexts and used to explore which factors influence durations most. Major…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+2

Speaking style adaptation in Text-To-Speech synthesis using Sequence-to-sequence models with attention

2018-10-29 · Bajibabu Bollepalli, Lauri Juvela, Paavo Alku

Currently, there are increasing interests in text-to-speech (TTS) synthesis to use sequence-to-sequence models with attention. These models are end-to-end meaning that they learn both co-articulation and duration propert…

Speech Synthesistext-to-speechText to SpeechText-To-Speech Synthesis+1