paper-with-me

홈 › Papers

Uncertainty Estimation in Deep Speech Enhancement Using Complex Gaussian Mixture Models

2022-12-09 · Huajian Fang, Timo Gerkmann

Single-channel deep speech enhancement approaches often estimate a single multiplicative mask to extract clean speech without a measure of its accuracy. Instead, in this work, we propose to quantify the uncertainty associated with clean speech estimates in neural network-based speech enhancement. Predictive uncertainty is typically categorized into aleatoric uncertainty and epistemic uncertainty. The former accounts for the inherent uncertainty in data and the latter corresponds to the model uncertainty. Aiming for robust clean speech estimation and efficient predictive uncertainty quantification, we propose to integrate statistical complex Gaussian mixture models (CGMMs) into a deep speech enhancement framework. More specifically, we model the dependency between input and output stochastically by means of a conditional probability density and train a neural network to map the noisy input to the full posterior distribution of clean speech, modeled as a mixture of multiple complex Gaussian components. Experimental results on different datasets show that the proposed algorithm effectively captures predictive uncertainty and that combining powerful statistical models and deep learning also delivers a superior speech enhancement performance.

📄 PDF Abstract BibTeX arXiv:2212.04831

Code (0)

등록된 구현이 없습니다.

Tasks

Speech EnhancementUncertainty Quantification

Similar Papers 제목 키워드 기반

Leveraging Heteroscedastic Uncertainty in Learning Complex Spectral Mapping for Single-channel Speech Enhancement

2022-11-16 · Kuan-Lin Chen, Daniel D. E. Wong, Ke Tan, Buye Xu 외

Most speech enhancement (SE) models learn a point estimate and do not make use of uncertainty estimation in the learning process. In this paper, we show that modeling heteroscedastic uncertainty by minimizing a multivari…

Speech Enhancement

Integrating Statistical Uncertainty into Neural Network-Based Speech Enhancement

2022-03-04 · Huajian Fang, Tal Peer, Stefan Wermter, Timo Gerkmann

Speech enhancement in the time-frequency domain is often performed by estimating a multiplicative mask to extract clean speech. However, most neural network-based methods perform point estimation, i.e., their output cons…

Speech Enhancement

Integrating Uncertainty into Neural Network-based Speech Enhancement

2023-05-15 · Huajian Fang, Dennis Becker, Stefan Wermter, Timo Gerkmann

Supervised masking approaches in the time-frequency domain aim to employ deep neural networks to estimate a multiplicative mask to extract clean speech. This leads to a single estimate for each input without any guarante…

Speech Enhancement

A weighted-variance variational autoencoder model for speech enhancement

2022-11-02 · Ali Golmakani, Mostafa Sadeghi, Xavier Alameda-Pineda, Romain Serizel

We address speech enhancement based on variational autoencoders, which involves learning a speech prior distribution in the time-frequency (TF) domain. A zero-mean complex-valued Gaussian distribution is usually assumed …

Speech Enhancement

Complex Recurrent Variational Autoencoder with Application to Speech Enhancement

2022-04-05 · Yuying Xie, Thomas Arildsen, Zheng-Hua Tan

As an extension of variational autoencoder (VAE), complex VAE uses complex Gaussian distributions to model latent variables and data. This work proposes a complex recurrent VAE framework, specifically in which complex-va…

Speech Enhancement