A Bayesian Permutation training deep representation learning method for speech enhancement with variational autoencoder
Recently, variational autoencoder (VAE), a deep representation learning (DRL) model, has been used to perform speech enhancement (SE). However, to the best of our knowledge, current VAE-based SE methods only apply VAE to the model speech signal, while noise is modeled using the traditional non-negative matrix factorization (NMF) model. One of the most important reasons for using NMF is that these VAE-based methods cannot disentangle the speech and noise latent variables from the observed signal. Based on Bayesian theory, this paper derives a novel variational lower bound for VAE, which ensures that VAE can be trained in supervision, and can disentangle speech and noise latent variables from the observed signal. This means that the proposed method can apply the VAE to model both speech and noise signals, which is totally different from the previous VAE-based SE works. More specifically, the proposed DRL method can learn to impose speech and noise signal priors to different sets of latent variables for SE. The experimental results show that the proposed method can not only disentangle speech and noise latent variables from the observed signal but also obtain a higher scale-invariant signal-to-distortion ratio and speech quality score than the similar deep neural network-based (DNN) SE method.
Code (0)
등록된 구현이 없습니다.
Tasks
Representation LearningSpeech EnhancementSimilar Papers 제목 키워드 기반
A deep representation learning speech enhancement method using $β$-VAE
In previous work, we proposed a variational autoencoder-based (VAE) Bayesian permutation training speech enhancement (SE) method (PVAE) which indicated that the SE performance of the traditional deep neural network-based…
DisentanglementRepresentation LearningSpeech EnhancementRobust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
This paper addresses the prevalent issue of incorrect speech output in audio-visual speech enhancement (AVSE) systems, which is often caused by poor video quality and mismatched training and test data. We introduce a pos…
Speech EnhancementPerceive and predict: self-supervised speech representation based loss functions for speech enhancement
Recent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the d…
Speech EnhancementA Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement
Self-supervised learned models have been found to be very effective for certain speech tasks such as automatic speech recognition, speaker identification, keyword spotting and others. While the features are undeniably us…
Automatic Speech RecognitionKeyword SpottingKnowledge DistillationSpeaker Identification+3Supervised and Unsupervised Speech Enhancement Using Nonnegative Matrix Factorization
Reducing the interference noise in a monaural noisy speech signal has been a challenging task for many years. Compared to traditional unsupervised speech enhancement methods, e.g., Wiener filtering, supervised approaches…
DenoisingSpeech DenoisingSpeech Enhancement