CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement
In this work, we further develop the conformer-based metric generative adversarial network (CMGAN) model for speech enhancement (SE) in the time-frequency (TF) domain. This paper builds on our previous work but takes a more in-depth look by conducting extensive ablation studies on model inputs and architectural design choices. We rigorously tested the generalization ability of the model to unseen noise types and distortions. We have fortified our claims through DNS-MOS measurements and listening tests. Rather than focusing exclusively on the speech denoising task, we extend this work to address the dereverberation and super-resolution tasks. This necessitated exploring various architectural changes, specifically metric discriminator scores and masking techniques. It is essential to highlight that this is among the earliest works that attempted complex TF-domain super-resolution. Our findings show that CMGAN outperforms existing state-of-the-art methods in the three major speech enhancement tasks: denoising, dereverberation, and super-resolution. For example, in the denoising task using the Voice Bank+DEMAND dataset, CMGAN notably exceeded the performance of prior models, attaining a PESQ score of 3.41 and an SSNR of 11.10 dB. Audio samples and CMGAN implementations are available online.
Code (2)
Tasks
Audio Super-ResolutionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDenoisingGenerative Adversarial NetworkSpeech DenoisingSpeech Enhancementspeech-recognitionSpeech RecognitionSpeech SeparationSuper-ResolutionSimilar Papers 제목 키워드 기반
CMGAN: Conformer-based Metric GAN for Speech Enhancement
Recently, convolution-augmented transformer (Conformer) has achieved promising performance in automatic speech recognition (ASR) and time-domain speech enhancement (SE), as it can capture both local and global dependenci…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderGenerative Adversarial Network+3Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement with Conformer-based Metric GAN
With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally ou…
Speech EnhancementPCNN: A Lightweight Parallel Conformer Neural Network for Efficient Monaural Speech Enhancement
Convolutional neural networks (CNN) and Transformer have wildly succeeded in multimedia applications. However, more effort needs to be made to harmonize these two architectures effectively to satisfy speech enhancement. …
Speech EnhancementSelective State Space Model for Monaural Speech Enhancement
Voice user interfaces (VUIs) have facilitated the efficient interactions between humans and machines through spoken commands. Since real-word acoustic scenes are complex, speech enhancement plays a critical role for robu…
MambaSpeech EnhancementSE Territory: Monaural Speech Enhancement Meets the Fixed Virtual Perceptual Space Mapping
Monaural speech enhancement has achieved remarkable progress recently. However, its performance has been constrained by the limited spatial cues available at a single microphone. To overcome this limitation, we introduce…
Multi-Task LearningSpeech Enhancement