Multi-reference Tacotron by Intercross Training for Style Disentangling,Transfer and Control in Speech Synthesis
Speech style control and transfer techniques aim to enrich the diversity and expressiveness of synthesized speech. Existing approaches model all speech styles into one representation, lacking the ability to control a specific speech feature independently. To address this issue, we introduce a novel multi-reference structure to Tacotron and propose intercross training approach, which together ensure that each sub-encoder of the multi-reference encoder independently disentangles and controls a specific style. Experimental results show that our model is able to control and transfer desired speech styles individually.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversitySpeech SynthesisMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Predicting Expressive Speaking Style From Text In End-To-End Speech Synthesis
Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-art end-to-end text-to-speech synthesis sy…
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisEnd-to-End Emotional Speech Synthesis Using Style Tokens and Semi-Supervised Training
This paper proposes an end-to-end emotional speech synthesis (ESS) method which adopts global style tokens (GSTs) for semi-supervised training. This model is built based on the GST-Tacotron framework. The style tokens ar…
Emotional Speech SynthesisEmotion RecognitionSpeech SynthesisExpressive Text-to-Speech using Style Tag
As recent text-to-speech (TTS) systems have been rapidly improved in speech quality and generation speed, many researchers now focus on a more challenging issue: expressive TTS. To control speaking styles, existing expre…
Language ModelingLanguage ModellingTAGtext-to-speech+1Whispered and Lombard Neural Speech Synthesis
It is desirable for a text-to-speech system to take into account the environment where synthetic speech is presented, and provide appropriate context-dependent output to the user. In this paper, we present and compare va…
Speaker VerificationSpeech Synthesistext-to-speechText to SpeechTowards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron
We present an extension to the Tacotron speech synthesis architecture that learns a latent embedding space of prosody, derived from a reference acoustic representation containing the desired prosody. We show that conditi…
Expressive Speech SynthesisSpeech Synthesis