Generative Adversarial Training for Text-to-Speech Synthesis Based on Raw Phonetic Input and Explicit Prosody Modelling
We describe an end-to-end speech synthesis system that uses generative adversarial training. We train our Vocoder for raw phoneme-to-audio conversion, using explicit phonetic, pitch and duration modeling. We experiment with several pre-trained models for contextualized and decontextualized word embeddings and we introduce a new method for highly expressive character voice matching, based on discreet style tokens.
Code (1)
Tasks
Speech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisWord EmbeddingsSimilar Papers 제목 키워드 기반
GANtron: Emotional Speech Synthesis with Generative Adversarial Networks
Speech synthesis is used in a wide variety of industries. Nonetheless, it always sounds flat or robotic. The state of the art methods that allow for prosody control are very cumbersome to use and do not allow easy tuning…
Emotional Speech SynthesisSpeech Synthesistext-to-speechText to SpeechDiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs
Denoising diffusion probabilistic models (DDPMs) are expressive generative models that have been used to solve a variety of speech synthesis problems. However, because of their high sampling costs, DDPMs are difficult to…
DenoisingSpeech Synthesistext-to-speechText to SpeechWaveform generation for text-to-speech synthesis using pitch-synchronous multi-scale generative adversarial networks
The state-of-the-art in text-to-speech synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference pro…
Image GenerationSpeech Synthesistext-to-speechText to Speech+2RSD-GAN: Regularized Sobolev Defense GAN Against Speech-to-Text Adversarial Attacks
This paper introduces a new synthesis-based defense algorithm for counteracting with a varieties of adversarial attacks developed for challenging the performance of the cutting-edge speech-to-text transcription systems. …
Speech-to-TextGenerative adversarial network-based glottal waveform model for statistical parametric speech synthesis
Recent studies have shown that text-to-speech synthesis quality can be improved by using glottal vocoding. This refers to vocoders that parameterize speech into two parts, the glottal excitation and vocal tract, that occ…
Generative Adversarial NetworkSpeech Synthesistext-to-speechText to Speech+1