Generative Adversarial Network based Speaker Adaptation for High Fidelity WaveNet Vocoder
Neural networks based vocoders, typically the WaveNet, have achieved spectacular performance for text-to-speech (TTS) in recent years. Although state-of-the-art parallel WaveNet has addressed the issue of real-time waveform generation, there remains problems. Firstly, due to the noisy input signal of the model, there is still a gap between the quality of generated and natural waveforms. Secondly, a parallel WaveNet is trained under a distilled training framework, which makes it tedious to adapt a well trained model to a new speaker. To address these two problems, this paper proposes an end-to-end adaptation method based on the generative adversarial network (GAN), which can reduce the computational cost for the training of new speaker adaptation. Our subjective experiments shows that the proposed training method can further reduce the quality gap between generated and natural waveforms.
Code (0)
등록된 구현이 없습니다.
Tasks
Generative Adversarial Networktext-to-speechText to SpeechVocal Bursts Intensity PredictionSimilar Papers 제목 키워드 기반
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Several recent work on speech synthesis have employed generative adversarial networks (GANs) to produce raw waveforms. Although such methods improve the sampling efficiency and memory usage, their sample quality has not …
CPUGPUSpeech SynthesisSpeakerGAN: Recognizing Speakers in New Languages with Generative Adversarial Networks
Verifying a person's identity based on their voice is a challenging, real-world problem in biometric security. A crucial requirement of such speaker verification systems is to be domain robust. Performance should not d…
Speaker VerificationGANSpeech: Adversarial Training for High-Fidelity Multi-Speaker Speech Synthesis
Recent advances in neural multi-speaker text-to-speech (TTS) models have enabled the generation of reasonably good speech quality with a single model and made it possible to synthesize the speech of a speaker with limite…
Speech Synthesistext-to-speechText to SpeechVocal Bursts Intensity PredictionDiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs
Denoising diffusion probabilistic models (DDPMs) are expressive generative models that have been used to solve a variety of speech synthesis problems. However, because of their high sampling costs, DDPMs are difficult to…
DenoisingSpeech Synthesistext-to-speechText to SpeechGenerative Adversarial Training Data Adaptation for Very Low-resource Automatic Speech Recognition
It is important to transcribe and archive speech data of endangered languages for preserving heritages of verbal culture and automatic speech recognition (ASR) is a powerful tool to facilitate this process. However, sinc…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cultural Vocal Bursts Intensity Predictionspeech-recognition+2