paper-with-me

홈 › Papers

Exploring synthetic data for cross-speaker style transfer in style representation based TTS

2024-09-25 · Lucas H. Ueda, Leonardo B. de M. M. Marques, Flávio O. Simões, Mário U. Neto, Fernando Runstein, Bianca Dal Bó, Paula D. P. Costa

Incorporating cross-speaker style transfer in text-to-speech (TTS) models is challenging due to the need to disentangle speaker and style information in audio. In low-resource expressive data scenarios, voice conversion (VC) can generate expressive speech for target speakers, which can then be used to train the TTS model. However, the quality and style transfer ability of the VC model are crucial for the overall TTS model quality. In this work, we explore the use of synthetic data generated by a VC model to assist the TTS model in cross-speaker style transfer tasks. Additionally, we employ pre-training of the style encoder using timbre perturbation and prototypical angular loss to mitigate speaker leakage. Our results show that using VC synthetic data can improve the naturalness and speaker similarity of TTS in cross-speaker scenarios. Furthermore, we extend this approach to a cross-language scenario, enhancing accent transfer.

📄 PDF Abstract BibTeX arXiv:2409.17364

Code (0)

등록된 구현이 없습니다.

Tasks

Style Transfertext-to-speechText to SpeechVoice Conversion

Similar Papers 제목 키워드 기반

Multi-speaker Multi-style Text-to-speech Synthesis With Single-speaker Single-style Training Data Scenarios

2021-12-23 · Qicong Xie, Tao Li, Xinsheng Wang, Zhichao Wang 외

In the existing cross-speaker style transfer task, a source speaker with multi-style recordings is necessary to provide the style for a target speaker. However, it is hard for one speaker to express all expected styles. …

DiversitySpeech SynthesisStyle Transfertext-to-speech+2

Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

2022-07-13 · Yookyung Shin, Younggun Lee, Suhee Jo, Yeongtae Hwang 외

Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the targe…

Language ModelingLanguage ModellingStyle Transfertext-to-speech+1

Cross-speaker style transfer for text-to-speech using data augmentation

2022-02-10 · Manuel Sam Ribeiro, Julian Roth, Giulia Comini, Goeric Huybrechts 외

We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting…

Data AugmentationStyle Transfertext-to-speechText to Speech+1

DeepTalk: Vocal Style Encoding for Speaker Recognition and Speech Synthesis

2020-12-09 · Anurag Chowdhury, Arun Ross, Prabu David

Automatic speaker recognition algorithms typically characterize speech audio using short-term spectral features that encode the physiological and anatomical aspects of speech production. Such algorithms do not fully capi…

Speaker RecognitionSpeech Synthesis

Improving the quality of neural TTS using long-form content and multi-speaker multi-style modeling

2022-12-20 · Tuomo Raitio, Javier Latorre, Andrea Davis, Tuuli Morrill 외

Neural text-to-speech (TTS) can provide quality close to natural speech if an adequate amount of high-quality speech material is available for training. However, acquiring speech data for TTS training is costly and time-…

Formtext-to-speechText to Speech