A Real-Time Wideband Neural Vocoder at 1.6 kb/s Using LPCNet
Neural speech synthesis algorithms are a promising new approach for coding speech at very low bitrate. They have so far demonstrated quality that far exceeds traditional vocoders, at the cost of very high complexity. In this work, we present a low-bitrate neural vocoder based on the LPCNet model. The use of linear prediction and sparse recurrent networks makes it possible to achieve real-time operation on general-purpose hardware. We demonstrate that LPCNet operating at 1.6 kb/s achieves significantly higher quality than MELP and that uncompressed LPCNet can exceed the quality of a waveform codec operating at low bitrate. This opens the way for new codec designs based on neural synthesis models.
Code (2)
Tasks
Speech SynthesisSimilar Papers 제목 키워드 기반
A Streamwise GAN Vocoder for Wideband Speech Coding at Very Low Bit Rate
Recently, GAN vocoders have seen rapid progress in speech synthesis, starting to outperform autoregressive models in perceptual quality with much higher generation speed. However, autoregressive vocoders are still the co…
Speech SynthesisBunched LPCNet : Vocoder for Low-cost Neural Text-To-Speech Systems
LPCNet is an efficient vocoder that combines linear prediction and deep neural network modules to keep the computational complexity low. In this work, we present two techniques to further reduce it's complexity, aiming f…
text-to-speechText to SpeechBunched LPCNet2: Efficient Neural Vocoders Covering Devices from Cloud to Edge
Text-to-Speech (TTS) services that run on edge devices have many advantages compared to cloud TTS, e.g., latency and privacy issues. However, neural vocoders with a low complexity and small model footprint inevitably gen…
Computational Efficiencytext-to-speechText to SpeechEnd-to-end LPCNet: A Neural Vocoder With Fully-Differentiable LPC Estimation
Neural vocoders have recently demonstrated high quality speech synthesis, but typically require a high computational complexity. LPCNet was proposed as a way to reduce the complexity of neural synthesis by using linear p…
Speech SynthesisControllable Sequence-To-Sequence Neural TTS with LPCNET Backend for Real-time Speech Synthesis on CPU
State-of-the-art sequence-to-sequence acoustic networks, that convert a phonetic sequence to a sequence of spectral features with no explicit prosody prediction, generate speech with close to natural quality, when cascad…
CPUProsody PredictionSentenceSpeech Synthesis