paper-with-me

Papers

Towards speech enhancement using a variational U-Net architecture

2020-12-07 · Eike J. Nustede, Jörn Anemüller

We investigate the viability of a variational U-Net architecture for denoising of single-channel audio data. Deep network speech enhancement systems commonly aim to estimate filter masks, or opt to work on the waveform signal, potentially neglecting relationships across higher dimensional spectro-temporal features. We study the adoption of a probabilistic bottleneck into the classic U-Net architecture for direct spectral reconstruction. Evaluation of several ablation network variants is carried out using signal-to-distortion ratio and perceptual measures, on audio data that includes known and unknown noise types as well as reverberation. Our experiments show that the residual (skip) connections in the proposed system are a prerequisite for successful spectral reconstruction, i.e., without filter mask estimation. Results show, on average, an advantage of the proposed variational U-Net architecture over its classic, non-variational version in signal enhancement performance under reverberant conditions of 0.31 and 6.98 in PESQ and STOI scores, respectively. Anecdotal evidence points to improved suppression of impulsive noise sources with the variational U-Net compared to the recurrent mask estimation network baseline.

📄 PDF Abstract BibTeX arXiv:2012.03594

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpectral ReconstructionSpeech Enhancement

Methods 이 논문이 사용한 방법론

Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

A Recurrent Variational Autoencoder for Speech Enhancement

2019-10-24 · Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin, Radu Horaud

This paper presents a generative approach to speech enhancement based on a recurrent variational autoencoder (RVAE). The deep generative speech model is trained using clean speech signals only, and it is combined with a …

Speech Enhancement

Switching Variational Auto-Encoders for Noise-Agnostic Audio-visual Speech Enhancement

2021-02-08 · Mostafa Sadeghi, Xavier Alameda-Pineda

Recently, audio-visual speech enhancement has been tackled in the unsupervised settings based on variational auto-encoders (VAEs), where during training only clean data is used to train a generative model for speech, whi…

Speech Enhancement

Disentanglement Learning for Variational Autoencoders Applied to Audio-Visual Speech Enhancement

2021-05-19 · Guillaume Carbajal, Julius Richter, Timo Gerkmann

Recently, the standard variational autoencoder has been successfully used to learn a probabilistic prior over speech signals, which is then used to perform speech enhancement. Variational autoencoders have then been cond…

AttributeDecoderDisentanglementSpeech Enhancement

Mixture of Inference Networks for VAE-based Audio-visual Speech Enhancement

2019-12-23 · Mostafa Sadeghi, Xavier Alameda-Pineda

In this paper, we are interested in unsupervised (unknown noise) audio-visual speech enhancement based on variational autoencoders (VAEs), where the probability distribution of clean speech spectra is simulated using an …

DecoderSpeech EnhancementVariational Inference

Unsupervised Speech Enhancement using Dynamical Variational Auto-Encoders

2021-06-23 · Xiaoyu Bie, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin

Dynamical variational autoencoders (DVAEs) are a class of deep generative models with latent variables, dedicated to model time series of high-dimensional data. DVAEs can be considered as extensions of the variational au…

Representation LearningSpeech EnhancementTime SeriesTime Series Analysis