DEEP COMPLEX-VALUED NEURAL BEAMFORMERS
We propose a complex-valued deep neural network (cDNN) for speech enhancement and source separation. While existing end-to-end systems use complex-valued gradients to pass the training error to a real-valued DNN used for gain mask estimation, we use the full potential of complex-valued LSTMs, MLPs and activation functions to estimate complex-valued beamforming weights directly from complex-valued microphone array data. By doing so, our cDNN is able to locate and track different moving sources by exploiting the phase information in the data. In our experiments, we use a typical living room environment, mixtures of the WallStreet Journal corpus, and YouTube noise. We compare our cDNN against the BeamformIt toolkit as a baseline, and a mask-based beamformer as a state-of-the-art reference system. We observed a significant improvement in terms of PESQ, STOI and WER.
Code (1)
Tasks
Speech EnhancementSimilar Papers 제목 키워드 기반
A Time-domain Real-valued Generalized Wiener Filter for Multi-channel Neural Separation Systems
Frequency-domain beamformers have been successful in a wide range of multi-channel neural separation systems in the past years. However, the operations in conventional frequency-domain beamformers are typically independe…
Speech SeparationSmoothed SVD-based Beamforming for FBMC/OQAM Systems Based on Frequency Spreading
The combination of singular value decomposition (SVD)-based beamforming and filter bank multicarrier with offset quadrature amplitude modulation (FBMC/OQAM) has not been successful to date. The difficulty of this combina…
A Low-complexity Structured Neural Network Approach to Intelligently Realize Wideband Multi-beam Beamformers
True-time-delay (TTD) beamformers can produce wideband, squint-free beams in both analog and digital signal domains, unlike frequency-dependent FFT beams. Our previous work showed that TTD beamformers can be efficiently …
Dynamic Independent Component/Vector Analysis: Time-Variant Linear Mixtures Separable by Time-Invariant Beamformers
A novel extension of Independent Component and Independent Vector Analysis for blind extraction/separation of one or several sources from time-varying mixtures is proposed. The mixtures are assumed to be separable source…
Localizing Spatial Information in Neural Spatiospectral Filters
Beamforming for multichannel speech enhancement relies on the estimation of spatial characteristics of the acoustic scene. In its simplest form, the delay-and-sum beamformer (DSB) introduces a time delay to all channels …
Speech Enhancement