Complex-valued Spatial Autoencoders for Multichannel Speech Enhancement
In this contribution, we present a novel online approach to multichannel speech enhancement. The proposed method estimates the enhanced signal through a filter-and-sum framework. More specifically, complex-valued masks are estimated by a deep complex-valued neural network, termed the complex-valued spatial autoencoder. The proposed network is capable of exploiting as well as manipulating both the phase and the amplitude of the microphone signals. As shown by the experimental results, the proposed approach is able to exploit both spatial and spectral characteristics of the desired source signal resulting in a physically plausible spatial selectivity and superior speech quality compared to other baseline methods.
Code (1)
Tasks
Speech EnhancementSimilar Papers 제목 키워드 기반
Spatially constrained vs. unconstrained filtering in neural spatiospectral filters for multichannel speech enhancement
When using artificial neural networks for multichannel speech enhancement, filtering is often achieved by estimating a complex-valued mask that is applied to all or one reference channel of the input signal. The estimati…
Speech EnhancementLocalizing Spatial Information in Neural Spatiospectral Filters
Beamforming for multichannel speech enhancement relies on the estimation of spatial characteristics of the acoustic scene. In its simplest form, the delay-and-sum beamformer (DSB) introduces a time delay to all channels …
Speech EnhancementSemi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization
In this paper we address speaker-independent multichannel speech enhancement in unknown noisy environments. Our work is based on a well-established multichannel local Gaussian modeling framework. We propose to use a neur…
Speech EnhancementLeveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the tempo…
MambaSpeech EnhancementSemi-Supervised Multichannel Speech Enhancement With a Deep Speech Prior
This paper describes a semi-supervised multichannel speech enhancement method that uses clean speech data for prior training. Although multichannel nonnegative matrix factorization (MNMF) and its constrained variant call…
Speech Enhancement