paper-with-me

홈 › Papers

Deep learning for minimum mean-square error approaches to speech enhancement

2019-08-01 · Speech communication 2019 8 · Aaron Nicolson, Kuldip K. Paliwal

Recently, the focus of speech enhancement research has shifted from minimum mean-square error (MMSE) approaches, like the MMSE short-time spectral amplitude (MMSE-STSA) estimator, to state-of-the-art masking- and mapping-based deep learning approaches. We aim to bridge the gap between these two differing speech enhancement approaches. Deep learning methods for MMSE approaches are investigated in this work, with the objective of producing intelligible enhanced speech at a high quality. Since the speech enhancement performance of an MMSE approach improves with the accuracy of the used a priori signal-to-noise ratio (SNR) estimator, a residual long short-term memory (ResLSTM) network is utilised here to accurately estimate the a priori SNR. MMSE approaches utilising the ResLSTM a priori SNR estimator are evaluated using subjective and objective measures of speech quality and intelligibility. The tested conditions include real-world non-stationary and coloured noise sources at multiple SNR levels. MMSE approaches utilising the proposed a priori SNR estimator are able to achieve higher enhanced speech quality and intelligibility scores than recent masking- and mapping-based deep learning approaches. The results presented in this work show that the performance of an MMSE approach to speech enhancement significantly increases when utilising deep learning. Availability: The proposed a priori SNR estimator is available at: https://github.com/anicolson/DeepXi.

📄 PDF Abstract BibTeX

Code (1)

anicolson/DeepXi tf

Tasks

Deep LearningSpeech Enhancement

Similar Papers 제목 키워드 기반

Beamforming design for minimizing the signal power estimation error

2025-06-20 · Esa Ollila, Xavier Mestre, Elias Raninen

We study the properties of beamformers in their ability to either maintain or estimate the true signal power of the signal of interest (SOI). Our focus is particularly on the Capon beamformer and the minimum mean squared…

Spoken Speech Enhancement using EEG

2019-09-13 · Gautam Krishna, Co Tran, Yan Han, Mason Carnahan 외

In this paper we demonstrate spoken speech enhancement using electroencephalography (EEG) signals using a generative adversarial network (GAN) based model, gated recurrent unit (GRU) regression based model, temporal conv…

EEGElectroencephalogram (EEG)Generative Adversarial Networkregression+1

Single-Channel Multi-talker Speech Recognition with Permutation Invariant Training

2017-07-19 · Yanmin Qian, Xuankai Chang, Dong Yu

Although great progresses have been made in automatic speech recognition (ASR), significant performance degradation is still observed when recognizing multi-talker mixed speech. In this paper, we propose and evaluate sev…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition+1

AV Speech Enhancement Challenge using a Real Noisy Corpus

2019-09-30 · Mandar Gogate, Ahsan Adeel, Kia Dashtipour, Peter Derleth 외

This paper presents, a first of its kind, audio-visual (AV) speech enhacement challenge in real-noisy settings. A detailed description of the AV challenge, a novel real noisy AV corpus (ASPIRE), benchmark speech enhancem…

Speech Enhancement

Data-Driven Prediction with Stochastic Data: Confidence Regions and Minimum Mean-Squared Error Estimates

2021-11-08 · Mingzhou Yin, Andrea Iannelli, Roy S. Smith

Recently, direct data-driven prediction has found important applications for controlling unknown systems, particularly in predictive control. Such an approach provides exact prediction using behavioral system theory when…

Predictionvalid