paper-with-me

Papers

A Streamwise GAN Vocoder for Wideband Speech Coding at Very Low Bit Rate

2021-08-09 · Ahmed Mustafa, Jan Büthe, Srikanth Korse, Kishan Gupta, Guillaume Fuchs, Nicola Pia

Recently, GAN vocoders have seen rapid progress in speech synthesis, starting to outperform autoregressive models in perceptual quality with much higher generation speed. However, autoregressive vocoders are still the common choice for neural generation of speech signals coded at very low bit rates. In this paper, we present a GAN vocoder which is able to generate wideband speech waveforms from parameters coded at 1.6 kbit/s. The proposed model is a modified version of the StyleMelGAN vocoder that can run in frame-by-frame manner, making it suitable for streaming applications. The experimental results show that the proposed model significantly outperforms prior autoregressive vocoders like LPCNet for very low bit rate speech coding, with computational complexity of about 5 GMACs, providing a new state of the art in this domain. Moreover, this streamwise adversarial vocoder delivers quality competitive to advanced speech codecs such as EVS at 5.9 kbit/s on clean speech, which motivates further usage of feed-forward fully-convolutional models for low bit rate speech coding.

📄 PDF Abstract BibTeX arXiv:2108.04051

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Synthesis

Similar Papers 제목 키워드 기반

NESC: Robust Neural End-2-End Speech Coding with GANs

2022-07-07 · Nicola Pia, Kishan Gupta, Srikanth Korse, Markus Multrus 외

Neural networks have proven to be a formidable tool to tackle the problem of speech coding at very low bit rates. However, the design of a neural coder that can be operated robustly under real-world conditions remains a …

Decoder

A Real-Time Wideband Neural Vocoder at 1.6 kb/s Using LPCNet

2019-03-28 · Jean-Marc Valin, Jan Skoglund

Neural speech synthesis algorithms are a promising new approach for coding speech at very low bitrate. They have so far demonstrated quality that far exceeds traditional vocoders, at the cost of very high complexity. In …

Speech Synthesis

Analysis by Adversarial Synthesis -- A Novel Approach for Speech Vocoding

2019-07-01 · Ahmed Mustafa, Arijit Biswas, Christian Bergler, Julia Schottenhamml 외

Classical parametric speech coding techniques provide a compact representation for speech signals. This affords a very low transmission rate but with a reduced perceptual quality of the reconstructed signals. Recently, a…

Unsupervised acoustic unit discovery for speech synthesis using discrete latent-variable neural networks

2019-04-16 · Ryan Eloff, André Nortje, Benjamin van Niekerk, Avashna Govender 외

For our submission to the ZeroSpeech 2019 challenge, we apply discrete latent-variable neural networks to unlabelled speech and use the discovered units for speech synthesis. Unsupervised discrete subword modelling could…

Acoustic Unit DiscoveryDecoderSpeech Synthesis

WARP-Q: Quality Prediction For Generative Neural Speech Codecs

2021-02-20 · Wissam A. Jassim, Jan Skoglund, Michael Chinen, Andrew Hines

Good speech quality has been achieved using waveform matching and parametric reconstruction coders. Recently developed very low bit rate generative codecs can reconstruct high quality wideband speech with bit streams les…

Dynamic Time WarpingPrediction