paper-with-me

Papers

Investigating Generative Adversarial Networks based Speech Dereverberation for Robust Speech Recognition

2018-03-27 · Ke Wang, Junbo Zhang, Sining Sun, Yujun Wang, Fei Xiang, Lei Xie

We investigate the use of generative adversarial networks (GANs) in speech dereverberation for robust speech recognition. GANs have been recently studied for speech enhancement to remove additive noises, but there still lacks of a work to examine their ability in speech dereverberation and the advantages of using GANs have not been fully established. In this paper, we provide deep investigations in the use of GAN-based dereverberation front-end in ASR. First, we study the effectiveness of different dereverberation networks (the generator in GAN) and find that LSTM leads a significant improvement as compared with feed-forward DNN and CNN in our dataset. Second, further adding residual connections in the deep LSTMs can boost the performance as well. Finally, we find that, for the success of GAN, it is important to update the generator and the discriminator using the same mini-batch data during training. Moreover, using reverberant spectrogram as a condition to discriminator, as suggested in previous studies, may degrade the performance. In summary, our GAN-based dereverberation front-end achieves 14%-19% relative CER reduction as compared to the baseline DNN dereverberation network when tested on a strong multi-condition training acoustic model.

📄 PDF Abstract BibTeX arXiv:1803.10132

Code (1)

wangkenpu/rsrgan 공식 구현 tf

Tasks

Robust Speech RecognitionSpeech DereverberationSpeech Enhancementspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Single-channel Speech Dereverberation via Generative Adversarial Training

2018-06-25 · Chenxing Li, Tieqiang Wang, Shuang Xu, Bo Xu

In this paper, we propose a single-channel speech dereverberation system (DeReGAT) based on convolutional, bidirectional long short-term memory and deep feed-forward neural network (CBLDNN) with generative adversarial tr…

Speech Dereverberation

RVAE-EM: Generative speech dereverberation based on recurrent variational auto-encoder and convolutive transfer function

2023-09-15 · Pengyu Wang, Xiaofei Li

In indoor scenes, reverberation is a crucial factor in degrading the perceived quality and intelligibility of speech. In this work, we propose a generative dereverberation method. Our approach is based on a probabilistic…

Speech Dereverberation

Analysing Diffusion-based Generative Approaches versus Discriminative Approaches for Speech Restoration

2022-11-04 · Jean-Marie Lemercier, Julius Richter, Simon Welker, Timo Gerkmann

Diffusion-based generative models have had a high impact on the computer vision and speech processing communities these past years. Besides data generation tasks, they have also been employed for data restoration tasks l…

Bandwidth ExtensionSpeech DenoisingSpeech DereverberationSpeech Enhancement

Schrödinger Bridge for Generative Speech Enhancement

2024-07-22 · Ante Jukić, Roman Korostik, Jagadeesh Balam, Boris Ginsburg

This paper proposes a generative speech enhancement model based on Schr\"odinger bridge (SB). The proposed model is employing a tractable SB to formulate a data-to-data process between the clean speech distribution and t…

DenoisingSpeech DenoisingSpeech DereverberationSpeech Enhancement

EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation

2024-06-10 · Julius Richter, Yi-Chiao Wu, Steven Krenn, Simon Welker 외

We release the EARS (Expressive Anechoic Recordings of Speech) dataset, a high-quality speech dataset comprising 107 speakers from diverse backgrounds, totaling in 100 hours of clean, anechoic speech data. The dataset co…

Speech Enhancement