paper-with-me

홈 › Papers

spatial-dccrn: dccrn equipped with frame-level angle feature and hybrid filtering for multi-channel speech enhancement

2022-10-17 · Shubo Lv, Yihui Fu, Yukai Jv, Lei Xie, Weixin Zhu, Wei Rao, Yannan Wang

Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network based masking estimation, we propose a multi-channel denoising neural network -- Spatial DCCRN. Firstly, we extend S-DCCRN to multi-channel scenario, aiming at performing cascaded sub-channel and full-channel processing strategy, which can model different channels separately. Moreover, instead of only adopting multi-channel spectrum or concatenating first-channel's magnitude and IPD as the model's inputs, we apply an angle feature extraction module (AFE) to extract frame-level angle feature embeddings, which can help the model to apparently perceive spatial information. Finally, since the phenomenon of residual noise will be more serious when the noise and speech exist in the same time frequency (TF) bin, we particularly design a masking and mapping filtering method to substitute the traditional filter-and-sum operation, with the purpose of cascading coarsely denoising, dereverberation and residual noise suppression. The proposed model, Spatial-DCCRN, has surpassed EaBNet, FasNet as well as several competitive models on the L3DAS22 Challenge dataset. Not only the 3D scenario, Spatial-DCCRN outperforms state-of-the-art (SOTA) model MIMO-UNet by a large margin in multiple evaluation metrics on the multi-channel ConferencingSpeech2021 Challenge dataset. Ablation studies also demonstrate the effectiveness of different contributions.

📄 PDF Abstract BibTeX arXiv:2210.08802

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingSpeech Enhancement

Similar Papers 제목 키워드 기반

A Deep Representation Learning-based Speech Enhancement Method Using Complex Convolution Recurrent Variational Autoencoder

2023-12-15 · Yang Xiang, Jingguang Tian, Xinhui Hu, Xinkang Xu 외

Generally, the performance of deep neural networks (DNNs) heavily depends on the quality of data representation learning. Our preliminary work has emphasized the significance of deep representation learning (DRL) in the …

Representation LearningSpeech Enhancement

DCCRN+: Channel-wise Subband DCCRN with SNR Estimation for Speech Enhancement

2021-06-16 · Shubo Lv, Yanxin Hu, Shimin Zhang, Lei Xie

Deep complex convolution recurrent network (DCCRN), which extends CRN with complex structure, has achieved superior performance in MOS evaluation in Interspeech 2020 deep noise suppression challenge (DNS2020). This paper…

DecoderSpeech Enhancement

DCCRN-KWS: an audio bias based model for noise robust small-footprint keyword spotting

2023-05-21 · Shubo Lv, Xiong Wang, Sining Sun, Long Ma 외

Real-world complex acoustic environments especially the ones with a low signal-to-noise ratio (SNR) will bring tremendous challenges to a keyword spotting (KWS) system. Inspired by the recent advances of neural speech en…

DenoisingKeyword SpottingMulti-Task LearningSmall-Footprint Keyword Spotting+3

S-DCCRN: Super Wide Band DCCRN with learnable complex feature for speech enhancement

2021-11-16 · Shubo Lv, Yihui Fu, Mengtao Xing, Jiayao Sun 외

In speech enhancement, complex neural network has shown promising performance due to their effectiveness in processing complex-valued spectrum. Most of the recent speech enhancement approaches mainly focus on wide-band s…

16kDenoisingSpeech DenoisingSpeech Enhancement

Multi-channel end-to-end neural network for speech enhancement, source localization, and voice activity detection

2022-06-20 · Yuan Chen, Yicheng Hsu, Mingsian R. Bai

Speech enhancement and source localization has been active research for several decades with a wide range of real-world applications. Recently, the Deep Complex Convolution Recurrent network (DCCRN) has yielded impressiv…

Action DetectionActivity DetectionSpeech Enhancement