paper-with-me

홈 › Papers

Improved MVDR Beamforming Using LSTM Speech Models to Clean Spatial Clustering Masks

2020-12-02 · Zhaoheng Ni, Felix Grezes, Viet Anh Trinh, Michael I. Mandel

Spatial clustering techniques can achieve significant multi-channel noise reduction across relatively arbitrary microphone configurations, but have difficulty incorporating a detailed speech/noise model. In contrast, LSTM neural networks have successfully been trained to recognize speech from noise on single-channel inputs, but have difficulty taking full advantage of the information in multi-channel recordings. This paper integrates these two approaches, training LSTM speech models to clean the masks generated by the Model-based EM Source Separation and Localization (MESSL) spatial clustering method. By doing so, it attains both the spatial separation performance and generality of multi-channel spatial clustering and the signal modeling performance of multiple parallel single-channel LSTM speech enhancers. Our experiments show that when our system is applied to the CHiME-3 dataset of noisy tablet recordings, it increases speech quality as measured by the Perceptual Evaluation of Speech Quality (PESQ) algorithm and reduces the word error rate of the baseline CHiME-3 speech recognizer, as compared to the default BeamformIt beamformer.

📄 PDF Abstract BibTeX arXiv:2012.02191

Code (0)

등록된 구현이 없습니다.

Tasks

Clustering

Methods 이 논문이 사용한 방법론

Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Subspace Hybrid Beamforming for Head-worn Microphone Arrays

2023-03-15 · Sina Hafezi, Alastair H. Moore, Pierre Guiraud, Patrick A. Naylor 외

A two-stage multi-channel speech enhancement method is proposed which consists of a novel adaptive beamformer, Hybrid Minimum Variance Distortionless Response (MVDR), Isotropic-MVDR (Iso), and a novel multi-channel spect…

DenoisingSpeech Enhancement

Multi-Talker MVDR Beamforming Based on Extended Complex Gaussian Mixture Model

2019-10-17

In this letter, we present a novel multi-talker minimum variance distortionless response (MVDR) beamforming as the front-end of an automatic speech recognition (ASR) system in a dinner party scenario. The CHiME-5 dataset…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+3

Beam-Guided TasNet: An Iterative Speech Separation Framework with Multi-Channel Output

2021-02-05 · Hangting Chen, Yang Yi, Dang Feng, Pengyuan Zhang

Time-domain audio separation network (TasNet) has achieved remarkable performance in blind source separation (BSS). Classic multi-channel speech processing framework employs signal estimation and beamforming. For example…

blind source separationSpeech Separation

All-neural beamformer for continuous speech separation

2021-10-13 · Zhuohuang Zhang, Takuya Yoshioka, Naoyuki Kanda, Zhuo Chen 외

Continuous speech separation (CSS) aims to separate overlapping voices from a continuous influx of conversational audio containing an unknown number of utterances spoken by an unknown number of speakers. A common applica…

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Extraction+3

Unsupervised Speech Enhancement Based on Multichannel NMF-Informed Beamforming for Noise-Robust Automatic Speech Recognition

2019-03-22 · Kazuki Shimada, Yoshiaki Bando, Masato Mimura, Katsutoshi Itoyama 외

This paper describes multichannel speech enhancement for improving automatic speech recognition (ASR) in noisy environments. Recently, the minimum variance distortionless response (MVDR) beamforming has widely been used …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Speech Enhancementspeech-recognition+1