paper-with-me

홈 › Papers

Generative Speech Coding with Predictive Variance Regularization

2021-02-18 · W. Bastiaan Kleijn, Andrew Storus, Michael Chinen, Tom Denton, Felicia S. C. Lim, Alejandro Luebs, Jan Skoglund, Hengchin Yeh

The recent emergence of machine-learning based generative models for speech suggests a significant reduction in bit rate for speech codecs is possible. However, the performance of generative models deteriorates significantly with the distortions present in real-world input signals. We argue that this deterioration is due to the sensitivity of the maximum likelihood criterion to outliers and the ineffectiveness of modeling a sum of independent signals with a single autoregressive model. We introduce predictive-variance regularization to reduce the sensitivity to outliers, resulting in a significant increase in performance. We show that noise reduction to remove unwanted signals can significantly increase performance. We provide extensive subjective performance evaluations that show that our system based on generative modeling provides state-of-the-art coding performance at 3 kb/s for real-world speech signals at reasonable computational complexity.

📄 PDF Abstract BibTeX arXiv:2102.09660

Code (1)

google/lyra

Tasks

Sensitivity

Similar Papers 제목 키워드 기반

Nonlinear predictive models computation in ADPCM schemes

2022-03-03 · Marcos Faundez-Zanuy

Recently several papers have been published on nonlinear prediction applied to speech coding. At ICASSP98 we presented a system based on an ADPCM scheme with a nonlinear predictor based on a neural net. The most critical…

Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

2023-05-18 · Hao Shi, Kazuki Shimada, Masato Hirano, Takashi Shibuya 외

Diffusion-based generative speech enhancement (SE) has recently received attention, but reverse diffusion remains time-consuming. One solution is to initialize the reverse diffusion process with enhanced features estimat…

DecoderSpeech Enhancement

A weighted-variance variational autoencoder model for speech enhancement

2022-11-02 · Ali Golmakani, Mostafa Sadeghi, Xavier Alameda-Pineda, Romain Serizel

We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed …

Speech Enhancement

Improved Speech Representations with Multi-Target Autoregressive Predictive Coding

2020-04-11 · ACL 2020 6 · Yu-An Chung, James Glass

Training objectives based on predictive coding have recently been shown to be very effective at learning meaningful representations from unlabeled speech. One example is Autoregressive Predictive Coding (Chung et al., 20…

speech-recognitionSpeech RecognitionTranslation

Regularizing Contrastive Predictive Coding for Speech Applications

2023-04-12 · Saurabhchand Bhati, Jesús Villalba, Piotr Żelasko, Laureano Moro-Velazquez 외

Self-supervised methods such as Contrastive predictive Coding (CPC) have greatly improved the quality of the unsupervised representations. These representations significantly reduce the amount of labeled data needed for …

Acoustic Unit DiscoveryAutomatic Speech RecognitionData Augmentationspeech-recognition+1