paper-with-me

홈 › Papers

Combining Spatial Clustering with LSTM Speech Models for Multichannel Speech Enhancement

2020-12-02 · Felix Grezes, Zhaoheng Ni, Viet Anh Trinh, Michael Mandel

Recurrent neural networks using the LSTM architecture can achieve significant single-channel noise reduction. It is not obvious, however, how to apply them to multi-channel inputs in a way that can generalize to new microphone configurations. In contrast, spatial clustering techniques can achieve such generalization, but lack a strong signal model. This paper combines the two approaches to attain both the spatial separation performance and generality of multichannel spatial clustering and the signal modeling performance of multiple parallel single-channel LSTM speech enhancers. The system is compared to several baselines on the CHiME3 dataset in terms of speech quality predicted by the PESQ algorithm and word error rate of a recognizer trained on mis-matched conditions, in order to focus on generalization. Our experiments show that by combining the LSTM models with the spatial clustering, we reduce word error rate by 4.6\% absolute (17.2\% relative) on the development set and 11.2\% absolute (25.5\% relative) on test set compared with spatial clustering system, and reduce by 10.75\% (32.72\% relative) on development set and 6.12\% absolute (15.76\% relative) on test data compared with LSTM model.

📄 PDF Abstract BibTeX arXiv:2012.03388

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringSpeech Enhancement

Methods 이 논문이 사용한 방법론

Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement

2024-09-16 · Wenze Ren, Haibin Wu, Yi-Cheng Lin, Xuanjun Chen 외

In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the tempo…

MambaSpeech Enhancement

Improved MVDR Beamforming Using LSTM Speech Models to Clean Spatial Clustering Masks

2020-12-02 · Zhaoheng Ni, Felix Grezes, Viet Anh Trinh, Michael I. Mandel

Spatial clustering techniques can achieve significant multi-channel noise reduction across relatively arbitrary microphone configurations, but have difficulty incorporating a detailed speech/noise model. In contrast, LST…

Clustering

Student-Teacher Learning for BLSTM Mask-based Speech Enhancement

2018-03-27

Spectral mask estimation using bidirectional long short-term memory (BLSTM) neural networks has been widely used in various speech enhancement applications, and it has achieved great success when it is applied to multich…

Speech Enhancementspeech-recognitionSpeech Recognition

Localizing Spatial Information in Neural Spatiospectral Filters

2023-03-14 · Annika Briegleb, Thomas Haubner, Vasileios Belagiannis, Walter Kellermann

Beamforming for multichannel speech enhancement relies on the estimation of spatial characteristics of the acoustic scene. In its simplest form, the delay-and-sum beamformer (DSB) introduces a time delay to all channels …

Speech Enhancement

Enhancement of Spatial Clustering-Based Time-Frequency Masks using LSTM Neural Networks

2020-12-02 · Felix Grezes, Zhaoheng Ni, Viet Anh Trinh, Michael Mandel

Recent works have shown that Deep Recurrent Neural Networks using the LSTM architecture can achieve strong single-channel speech enhancement by estimating time-frequency masks. However, these models do not naturally gene…

ClusteringSpeech Enhancement