paper-with-me

Papers

Multi-task WaveNet: A Multi-task Generative Model for Statistical Parametric Speech Synthesis without Fundamental Frequency Conditions

2018-06-22

This paper introduces an improved generative model for statistical parametric speech synthesis (SPSS) based on WaveNet under a multi-task learning framework. Different from the original WaveNet model, the proposed Multi-task WaveNet employs the frame-level acoustic feature prediction as the secondary task and the external fundamental frequency prediction model for the original WaveNet can be removed. Therefore the improved WaveNet can generate high-quality speech waveforms only conditioned on linguistic features. Multi-task WaveNet can produce more natural and expressive speech by addressing the pitch prediction error accumulation issue and possesses more succinct inference procedures than the original WaveNet. Experimental results prove that the SPSS method proposed in this paper can achieve better performance than the state-of-the-art approach utilizing the original WaveNet in both objective and subjective preference tests.

📄 PDF Abstract BibTeX arXiv:1806.08619

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task LearningPredictionSpeech Synthesis

Similar Papers 제목 키워드 기반

Stochastic WaveNet: A Generative Latent Variable Model for Sequential Data

2018-06-15 · Guokun Lai, Bohan Li, Guoqing Zheng, Yiming Yang

How to model distribution of sequential data, including but not limited to speech and human motions, is an important ongoing research problem. It has been demonstrated that model capacity can be significantly enhanced by…

Waveform generation for text-to-speech synthesis using pitch-synchronous multi-scale generative adversarial networks

2018-10-30 · Lauri Juvela, Bajibabu Bollepalli, Junichi Yamagishi, Paavo Alku

The state-of-the-art in text-to-speech synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference pro…

Image GenerationSpeech Synthesistext-to-speechText to Speech+2

Speaker-independent raw waveform model for glottal excitation

2018-04-25 · Lauri Juvela, Vassilis Tsiaras, Bajibabu Bollepalli, Manu Airaksinen 외

Recent speech technology research has seen a growing interest in using WaveNets as statistical vocoders, i.e., generating speech waveforms from acoustic features. These models have been shown to improve the generated spe…

modelSpeech Synthesistext-to-speechText to Speech+2

Generative Adversarial Network based Speaker Adaptation for High Fidelity WaveNet Vocoder

2018-12-06 · Qiao Tian, Bing Yang, Jing Chen, Benlai Tang 외

Neural networks based vocoders, typically the WaveNet, have achieved spectacular performance for text-to-speech (TTS) in recent years. Although state-of-the-art parallel WaveNet has addressed the issue of real-time wavef…

Generative Adversarial Networktext-to-speechText to SpeechVocal Bursts Intensity Prediction

Wavenet based low rate speech coding

2017-12-01 · W. Bastiaan Kleijn, Felicia S. C. Lim, Alejandro Luebs, Jan Skoglund 외

Traditional parametric coding of speech facilitates low rate but provides poor reconstruction quality because of the inadequacy of the model used. We describe how a WaveNet generative speech model can be used to generate…

Bandwidth Extension