Towards speech enhancement using a variational U-Net architecture
We investigate the viability of a variational U-Net architecture for denoising of single-channel audio data. Deep network speech enhancement systems commonly aim to estimate filter masks, or opt to work on the waveform signal, potentially neglecting relationships across higher dimensional spectro-temporal features. We study the adoption of a probabilistic bottleneck into the classic U-Net architecture for direct spectral reconstruction. Evaluation of several ablation network variants is carried out using signal-to-distortion ratio and perceptual measures, on audio data that includes known and unknown noise types as well as reverberation. Our experiments show that the residual (skip) connections in the proposed system are a prerequisite for successful spectral reconstruction, i.e., without filter mask estimation. Results show, on average, an advantage of the proposed variational U-Net architecture over its classic, non-variational version in signal enhancement performance under reverberant conditions of 0.31 and 6.98 in PESQ and STOI scores, respectively. Anecdotal evidence points to improved suppression of impulsive noise sources with the variational U-Net compared to the recurrent mask estimation network baseline.
Code (0)
등록된 구현이 없습니다.
Tasks
DenoisingSpectral ReconstructionSpeech EnhancementMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Recurrent Variational Autoencoder for Speech Enhancement
This paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a …
Speech EnhancementSwitching Variational Auto-Encoders for Noise-Agnostic Audio-visual Speech Enhancement
Recently, audio-visual speech enhancement has been tackled in the unsupervised settings based on variational auto-encoders (VAEs), where during training only clean data is used to train a generative model for speech, whi…
Speech EnhancementDisentanglement Learning for Variational Autoencoders Applied to Audio-Visual Speech Enhancement
Recently, the standard variational autoencoder has been successfully used to learn a probabilistic prior over speech signals, which is then used to perform speech enhancement. Variational autoencoders have then been cond…
AttributeDecoderDisentanglementSpeech EnhancementMixture of Inference Networks for VAE-based Audio-visual Speech Enhancement
In this paper, we are interested in unsupervised (unknown noise) audio-visual speech enhancement based on variational autoencoders (VAEs), where the probability distribution of clean speech spectra is simulated using an …
DecoderSpeech EnhancementVariational InferenceUnsupervised Speech Enhancement using Dynamical Variational Auto-Encoders
Dynamical variational autoencoders (DVAEs) are a class of deep generative models with latent variables, dedicated to model time series of high-dimensional data. DVAEs can be considered as extensions of the variational au…
Representation LearningSpeech EnhancementTime SeriesTime Series Analysis