paper-with-me

홈 › Papers

CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement

2022-09-22 · Sherif Abdulatif, Ruizhe Cao, Bin Yang

In this work, we further develop the conformer-based metric generative adversarial network (CMGAN) model for speech enhancement (SE) in the time-frequency (TF) domain. This paper builds on our previous work but takes a more in-depth look by conducting extensive ablation studies on model inputs and architectural design choices. We rigorously tested the generalization ability of the model to unseen noise types and distortions. We have fortified our claims through DNS-MOS measurements and listening tests. Rather than focusing exclusively on the speech denoising task, we extend this work to address the dereverberation and super-resolution tasks. This necessitated exploring various architectural changes, specifically metric discriminator scores and masking techniques. It is essential to highlight that this is among the earliest works that attempted complex TF-domain super-resolution. Our findings show that CMGAN outperforms existing state-of-the-art methods in the three major speech enhancement tasks: denoising, dereverberation, and super-resolution. For example, in the denoising task using the Voice Bank+DEMAND dataset, CMGAN notably exceeded the performance of prior models, attaining a PESQ score of 3.41 and an SSNR of 11.10 dB. Audio samples and CMGAN implementations are available online.

📄 PDF Abstract BibTeX arXiv:2209.11112

Code (2)

SherifAbdulatif/CMGAN 공식 구현 pytorch
ruizhecao96/cmgan 공식 구현 pytorch

Tasks

Audio Super-ResolutionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDenoisingGenerative Adversarial NetworkSpeech DenoisingSpeech Enhancementspeech-recognitionSpeech RecognitionSpeech SeparationSuper-Resolution

Similar Papers 제목 키워드 기반

CMGAN: Conformer-based Metric GAN for Speech Enhancement

2022-03-28 · Ruizhe Cao, Sherif Abdulatif, Bin Yang

Recently, convolution-augmented transformer (Conformer) has achieved promising performance in automatic speech recognition (ASR) and time-domain speech enhancement (SE), as it can capture both local and global dependenci…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderGenerative Adversarial Network+3

Unrestricted Global Phase Bias-Aware Single-channel Speech Enhancement with Conformer-based Metric GAN

2024-02-13 · Shiqi Zhang, Zheng Qiu, Daiki Takeuchi, Noboru Harada 외

With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally ou…

Speech Enhancement

PCNN: A Lightweight Parallel Conformer Neural Network for Efficient Monaural Speech Enhancement

2023-07-28 · Xinmeng Xu, Weiping tu, Yuhong Yang

Convolutional neural networks (CNN) and Transformer have wildly succeeded in multimedia applications. However, more effort needs to be made to harmonize these two architectures effectively to satisfy speech enhancement. …

Speech Enhancement

Selective State Space Model for Monaural Speech Enhancement

2024-11-09 · Moran Chen, Qiquan Zhang, Mingjiang Wang, Xiangyu Zhang 외

Voice user interfaces (VUIs) have facilitated the efficient interactions between humans and machines through spoken commands. Since real-word acoustic scenes are complex, speech enhancement plays a critical role for robu…

MambaSpeech Enhancement

SE Territory: Monaural Speech Enhancement Meets the Fixed Virtual Perceptual Space Mapping

2023-11-03 · Xinmeng Xu, Yuhong Yang, Weiping tu

Monaural speech enhancement has achieved remarkable progress recently. However, its performance has been constrained by the limited spatial cues available at a single microphone. To overcome this limitation, we introduce…

Multi-Task LearningSpeech Enhancement