Simultaneous Denoising and Dereverberation Using Deep Embedding Features
Monaural speech dereverberation is a very challenging task because no spatial cues can be used. When the additive noises exist, this task becomes more challenging. In this paper, we propose a joint training method for simultaneous speech denoising and dereverberation using deep embedding features, which is based on the deep clustering (DC). DC is a state-of-the-art method for speech separation that includes embedding learning and K-means clustering. As for our proposed method, it contains two stages: denoising and dereverberation. At the denoising stage, the DC network is leveraged to extract noise-free deep embedding features. These embedding features are generated from the anechoic speech and residual reverberation signals. They can represent the inferred spectral masking patterns of the desired signals, which are discriminative features. At the dereverberation stage, instead of using the unsupervised K-means clustering algorithm, another supervised neural network is utilized to estimate the anechoic speech from these deep embedding features. Finally, the denoising stage and dereverberation stage are optimized by the joint training method. Experimental results show that the proposed method outperforms the WPE and BLSTM baselines, especially in the low SNR condition.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringDeep ClusteringDenoisingSpeech DenoisingSpeech DereverberationSpeech SeparationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
D²Net: A Denoising and Dereverberation Network Based on Two-branch Encoder and Dual-path Transformer
The simultaneous denoising and dereverberation for single-channel mixture speech under the complicated acoustic environment is considered to be a challengeable task. In this paper, we propose a denoising and dereverberat…
DenoisingSpeech EnhancementBSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding
Speech self-supervised learning (SSL) represents has achieved state-of-the-art (SOTA) performance in multiple downstream tasks. However, its application in speech enhancement (SE) tasks remains immature, offering opportu…
DenoisingSelf-Supervised LearningSpeech EnhancementJointly optimal dereverberation and beamforming
We previously proposed an optimal (in the maximum likelihood sense) convolutional beamformer that can perform simultaneous denoising and dereverberation, and showed its superiority over the widely used cascade of a WPE d…
DenoisingReal-time Denoising and Dereverberation with Tiny Recurrent U-Net
Modern deep learning-based models have seen outstanding performance improvement with speech enhancement tasks. The number of parameters of state-of-the-art models, however, is often too large to be deployed on devices fo…
DenoisingSpeech EnhancementTo Dereverb Or Not to Dereverb? Perceptual Studies On Real-Time Dereverberation Targets
In real life, room effect, also known as room reverberation, and the present background noise degrade the quality of speech. Recently, deep learning-based speech enhancement approaches have shown a lot of promise and sur…
DenoisingSpeech Enhancement