paper-with-me

홈 › Papers

S-DCCRN: Super Wide Band DCCRN with learnable complex feature for speech enhancement

2021-11-16 · Shubo Lv, Yihui Fu, Mengtao Xing, Jiayao Sun, Lei Xie, Jun Huang, Yannan Wang, Tao Yu

In speech enhancement, complex neural network has shown promising performance due to their effectiveness in processing complex-valued spectrum. Most of the recent speech enhancement approaches mainly focus on wide-band signal with a sampling rate of 16K Hz. However, research on super wide band (e.g., 32K Hz) or even full-band (48K) denoising is still lacked due to the difficulty of modeling more frequency bands and particularly high frequency components. In this paper, we extend our previous deep complex convolution recurrent neural network (DCCRN) substantially to a super wide band version -- S-DCCRN, to perform speech denoising on speech of 32K Hz sampling rate. We first employ a cascaded sub-band and full-band processing module, which consists of two small-footprint DCCRNs -- one operates on sub-band signal and one operates on full-band signal, aiming at benefiting from both local and global frequency information. Moreover, instead of simply adopting the STFT feature as input, we use a complex feature encoder trained in an end-to-end manner to refine the information of different frequency bands. We also use a complex feature decoder to revert the feature to time-frequency domain. Finally, a learnable spectrum compression method is adopted to adjust the energy of different frequency bands, which is beneficial for neural network learning. The proposed model, S-DCCRN, has surpassed PercepNet as well as several competitive models and achieves state-of-the-art performance in terms of speech quality and intelligibility. Ablation studies also demonstrate the effectiveness of different contributions.

📄 PDF Abstract BibTeX arXiv:2111.08387

Code (0)

등록된 구현이 없습니다.

Tasks

16kDenoisingSpeech DenoisingSpeech Enhancement

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

DCCRN+: Channel-wise Subband DCCRN with SNR Estimation for Speech Enhancement

2021-06-16 · Shubo Lv, Yanxin Hu, Shimin Zhang, Lei Xie

Deep complex convolution recurrent network (DCCRN), which extends CRN with complex structure, has achieved superior performance in MOS evaluation in Interspeech 2020 deep noise suppression challenge (DNS2020). This paper…

DecoderSpeech Enhancement

A Deep Representation Learning-based Speech Enhancement Method Using Complex Convolution Recurrent Variational Autoencoder

2023-12-15 · Yang Xiang, Jingguang Tian, Xinhui Hu, Xinkang Xu 외

Generally, the performance of deep neural networks (DNNs) heavily depends on the quality of data representation learning. Our preliminary work has emphasized the significance of deep representation learning (DRL) in the …

Representation LearningSpeech Enhancement

Multi-channel end-to-end neural network for speech enhancement, source localization, and voice activity detection

2022-06-20 · Yuan Chen, Yicheng Hsu, Mingsian R. Bai

Speech enhancement and source localization has been active research for several decades with a wide range of real-world applications. Recently, the Deep Complex Convolution Recurrent network (DCCRN) has yielded impressiv…

Action DetectionActivity DetectionSpeech Enhancement

DCCRN-KWS: an audio bias based model for noise robust small-footprint keyword spotting

2023-05-21 · Shubo Lv, Xiong Wang, Sining Sun, Long Ma 외

Real-world complex acoustic environments especially the ones with a low signal-to-noise ratio (SNR) will bring tremendous challenges to a keyword spotting (KWS) system. Inspired by the recent advances of neural speech en…

DenoisingKeyword SpottingMulti-Task LearningSmall-Footprint Keyword Spotting+3

spatial-dccrn: dccrn equipped with frame-level angle feature and hybrid filtering for multi-channel speech enhancement

2022-10-17 · Shubo Lv, Yihui Fu, Yukai Jv, Lei Xie 외

Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network ba…

DenoisingSpeech Enhancement