Exploring synthetic data for cross-speaker style transfer in style representation based TTS
Incorporating cross-speaker style transfer in text-to-speech (TTS) models is challenging due to the need to disentangle speaker and style information in audio. In low-resource expressive data scenarios, voice conversion (VC) can generate expressive speech for target speakers, which can then be used to train the TTS model. However, the quality and style transfer ability of the VC model are crucial for the overall TTS model quality. In this work, we explore the use of synthetic data generated by a VC model to assist the TTS model in cross-speaker style transfer tasks. Additionally, we employ pre-training of the style encoder using timbre perturbation and prototypical angular loss to mitigate speaker leakage. Our results show that using VC synthetic data can improve the naturalness and speaker similarity of TTS in cross-speaker scenarios. Furthermore, we extend this approach to a cross-language scenario, enhancing accent transfer.
Code (0)
등록된 구현이 없습니다.
Tasks
Style Transfertext-to-speechText to SpeechVoice ConversionSimilar Papers 제목 키워드 기반
Multi-speaker Multi-style Text-to-speech Synthesis With Single-speaker Single-style Training Data Scenarios
In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker. However, it is hard for one speaker to express all expected styles. …
DiversitySpeech SynthesisStyle Transfertext-to-speech+2Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS
Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the targe…
Language ModelingLanguage ModellingStyle Transfertext-to-speech+1Cross-speaker style transfer for text-to-speech using data augmentation
We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting…
Data AugmentationStyle Transfertext-to-speechText to Speech+1DeepTalk: Vocal Style Encoding for Speaker Recognition and Speech Synthesis
Automatic speaker recognition algorithms typically characterize speech audio using short-term spectral features that encode the physiological and anatomical aspects of speech production. Such algorithms do not fully capi…
Speaker RecognitionSpeech SynthesisImproving the quality of neural TTS using long-form content and multi-speaker multi-style modeling
Neural text-to-speech (TTS) can provide quality close to natural speech if an adequate amount of high-quality speech material is available for training. However, acquiring speech data for TTS training is costly and time-…
Formtext-to-speechText to Speech