Towards Generalized Speech Enhancement with Generative Adversarial Networks
The speech enhancement task usually consists of removing additive noise or reverberation that partially mask spoken utterances, affecting their intelligibility. However, little attention is drawn to other, perhaps more aggressive signal distortions like clipping, chunk elimination, or frequency-band removal. Such distortions can have a large impact not only on intelligibility, but also on naturalness or even speaker identity, and require of careful signal reconstruction. In this work, we give full consideration to this generalized speech enhancement task, and show it can be tackled with a time-domain generative adversarial network (GAN). In particular, we extend a previous GAN-based speech enhancement system to deal with mixtures of four types of aggressive distortions. Firstly, we propose the addition of an adversarial acoustic regression loss that promotes a richer feature extraction at the discriminator. Secondly, we also make use of a two-step adversarial training schedule, acting as a warm up-and-fine-tune sequence. Both objective and subjective evaluations show that these two additions bring improved speech reconstructions that better match the original speaker identity and naturalness.
Code (0)
등록된 구현이 없습니다.
Tasks
Generative Adversarial NetworkSpeech EnhancementSimilar Papers 제목 키워드 기반
VSEGAN: Visual Speech Enhancement Generative Adversarial Network
Speech enhancement is an essential task of improving speech quality in noise scenario. Several state-of-the-art approaches have introduced visual information for speech enhancement,since the visual aspect of speech is es…
Generative Adversarial NetworkSpeech EnhancementTime-domain Speech Enhancement with Generative Adversarial Learning
Speech enhancement aims to obtain speech signals with high intelligibility and quality from noisy speech. Recent work has demonstrated the excellent performance of time-domain deep learning methods, such as Conv-TasNet. …
Generative Adversarial NetworkSpeech EnhancementiSEGAN: Improved Speech Enhancement Generative Adversarial Networks
Popular neural network-based speech enhancement systems operate on the magnitude spectrogram and ignore the phase mismatch between the noisy and clean speech signals. Conditional generative adversarial networks (cGANs) s…
Speech EnhancementAeGAN: Time-Frequency Speech Denoising via Generative Adversarial Networks
Automatic speech recognition (ASR) systems are of vital importance nowadays in commonplace tasks such as speech-to-text processing and language translation. This created the need for an ASR system that can operate in rea…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingGenerative Adversarial Network+6Conditional Diffusion Probabilistic Model for Speech Enhancement
Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models have shown strong potential in speech sy…
modelSpeech EnhancementSpeech Synthesis