paper-with-me

홈 › Papers

Enhancement and Recognition of Reverberant and Noisy Speech by Extending Its Coherence

2015-09-02 · Scott Wisdom, Thomas Powers, Les Atlas, James Pitton

Most speech enhancement algorithms make use of the short-time Fourier transform (STFT), which is a simple and flexible time-frequency decomposition that estimates the short-time spectrum of a signal. However, the duration of short STFT frames are inherently limited by the nonstationarity of speech signals. The main contribution of this paper is a demonstration of speech enhancement and automatic speech recognition in the presence of reverberation and noise by extending the length of analysis windows. We accomplish this extension by performing enhancement in the short-time fan-chirp transform (STFChT) domain, an overcomplete time-frequency representation that is coherent with speech signals over longer analysis window durations than the STFT. This extended coherence is gained by using a linear model of fundamental frequency variation of voiced speech signals. Our approach centers around using a single-channel minimum mean-square error log-spectral amplitude (MMSE-LSA) estimator proposed by Habets, which scales coefficients in a time-frequency domain to suppress noise and reverberation. In the case of multiple microphones, we preprocess the data with either a minimum variance distortionless response (MVDR) beamformer, or a delay-and-sum beamformer (DSB). We evaluate our algorithm on both speech enhancement and recognition tasks for the REVERB challenge dataset. Compared to the same processing done in the STFT domain, our approach achieves significant improvement in terms of objective enhancement metrics (including PESQ---the ITU-T standard measurement for speech quality). In terms of automatic speech recognition (ASR) performance as measured by word error rate (WER), our experiments indicate that the STFT with a long window is more effective for ASR.

📄 PDF Abstract BibTeX arXiv:1509.00533

Code (0)

등록된 구현이 없습니다.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Exploring Speech Enhancement with Generative Adversarial Networks for Robust Speech Recognition

2017-11-15 · Chris Donahue, Bo Li, Rohit Prabhavalkar

We investigate the effectiveness of generative adversarial networks (GANs) for speech enhancement, in the context of improving noise robustness of automatic speech recognition (ASR) systems. Prior work demonstrates that …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Robust Speech RecognitionSpeech Enhancement+2

Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement

2025-05-07 · Rauf Nasretdinov, Roman Korostik, Ante Jukić

In this work, we investigate application of generative speech enhancement to improve the robustness of ASR models in noisy and reverberant conditions. We employ a recently-proposed speech enhancement model based on Schr\…

Robust Speech RecognitionSpeech Enhancementspeech-recognitionSpeech Recognition

Brouhaha: multi-task training for voice activity detection, speech-to-noise ratio, and C50 room acoustics estimation

2022-10-24 · Marvin Lavechin, Marianne Métais, Hadrien Titeux, Alodie Boissonnet 외

Most automatic speech processing systems register degraded performance when applied to noisy or reverberant speech. But how can one tell whether speech is noisy or reverberant? We propose Brouhaha, a neural network joint…

Action DetectionActivity DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)+2

An Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch Conditions

2019-04-19 · Aswin Shanmugam Subramanian, Xiaofei Wang, Shinji Watanabe, Toru Taniguchi 외

Sequence-to-sequence (S2S) modeling is becoming a popular paradigm for automatic speech recognition (ASR) because of its ability to jointly optimize all the conventional ASR components in an end-to-end (E2E) fashion. Thi…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingNoisy Speech Recognition+3

Unsupervised Speech Enhancement with speech recognition embedding and disentanglement losses

2021-11-16 · Viet Anh Trinh, Sebastian Braun

Speech enhancement has recently achieved great success with various deep learning methods. However, most conventional speech enhancement systems are trained with supervised methods that impose two significant challenges.…

DisentanglementSpeech Enhancementspeech-recognitionSpeech Recognition