paper-with-me

홈 › Papers

Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition

2017-11-15 · Chris Donahue, Bo Li, Rohit Prabhavalkar

We investigate the effectiveness of generative adversarial networks (GANs) for speech enhancement, in the context of improving noise robustness of automatic speech recognition (ASR) systems. Prior work demonstrates that GANs can effectively suppress additive noise in raw waveform speech signals, improving perceptual quality metrics; however this technique was not justified in the context of ASR. In this work, we conduct a detailed study to measure the effectiveness of GANs in enhancing speech contaminated by both additive and reverberant noise. Motivated by recent advances in image processing, we propose operating GANs on log-Mel filterbank spectra instead of waveforms, which requires less computation and is more robust to reverberant noise. While GAN enhancement improves the performance of a clean-trained ASR system on noisy speech, it falls short of the performance achieved by conventional multi-style training (MTR). By appending the GAN-enhanced features to the noisy inputs and retraining, we achieve a 7% WER improvement relative to the MTR system.

📄 PDF Abstract BibTeX arXiv:1711.05747

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Robust Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

VSEGAN: Visual Speech Enhancement Generative Adversarial Network

2021-02-04 · Xinmeng Xu, Yang Wang, Dongxiang Xu, Yiyuan Peng 외

Speech enhancement is an essential task of improving speech quality in noise scenario. Several state-of-the-art approaches have introduced visual information for speech enhancement,since the visual aspect of speech is es…

Generative Adversarial NetworkSpeech Enhancement

AeGAN: Time-Frequency Speech Denoising via Generative Adversarial Networks

2019-10-21 · Sherif Abdulatif, Karim Armanious, Karim Guirguis, Jayasankar T. Sajeev 외

Automatic speech recognition (ASR) systems are of vital importance nowadays in commonplace tasks such as speech-to-text processing and language translation. This created the need for an ASR system that can operate in rea…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingGenerative Adversarial Network+6

iSEGAN: Improved Speech Enhancement Generative Adversarial Networks

2020-02-20 · Deepak Baby

Popular neural network-based speech enhancement systems operate on the magnitude spectrogram and ignore the phase mismatch between the noisy and clean speech signals. Conditional generative adversarial networks (cGANs) s…

Speech Enhancement

Unsupervised speech enhancement with diffusion-based generative models

2023-09-19 · Berné Nortier, Mostafa Sadeghi, Romain Serizel

Recently, conditional score-based diffusion models have gained significant attention in the field of supervised speech enhancement, yielding state-of-the-art performance. However, these methods may face challenges when g…

Speech Enhancement

Time-domain Speech Enhancement with Generative Adversarial Learning

2021-03-30 · Feiyang Xiao, Jian Guan, Qiuqiang Kong, Wenwu Wang

Speech enhancement aims to obtain speech signals with high intelligibility and quality from noisy speech. Recent work has demonstrated the excellent performance of time-domain deep learning methods, such as Conv-TasNet. …

Generative Adversarial NetworkSpeech Enhancement