paper-with-me

홈 › Papers

DCCRN+: Channel-wise Subband DCCRN with SNR Estimation for Speech Enhancement

2021-06-16 · Shubo Lv, Yanxin Hu, Shimin Zhang, Lei Xie

Deep complex convolution recurrent network (DCCRN), which extends CRN with complex structure, has achieved superior performance in MOS evaluation in Interspeech 2020 deep noise suppression challenge (DNS2020). This paper further extends DCCRN with the following significant revisions. We first extend the model to sub-band processing where the bands are split and merged by learnable neural network filters instead of engineered FIR filters, leading to a faster noise suppressor trained in an end-to-end manner. Then the LSTM is further substituted with a complex TF-LSTM to better model temporal dependencies along both time and frequency axes. Moreover, instead of simply concatenating the output of each encoder layer to the input of the corresponding decoder layer, we use convolution blocks to first aggregate essential information from the encoder output before feeding it to the decoder layers. We specifically formulate the decoder with an extra a priori SNR estimation module to maintain good speech quality while removing noise. Finally a post-processing module is adopted to further suppress the unnatural residual noise. The new model, named DCCRN+, has surpassed the original DCCRN as well as several competitive models in terms of PESQ and DNSMOS, and has achieved superior performance in the new Interspeech 2021 DNS challenge

📄 PDF Abstract BibTeX arXiv:2106.08672

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderSpeech Enhancement

Methods 이 논문이 사용한 방법론

CRN Conditional Relation Network, or CRN, is a building block to construct more sophisticated structures for representation and reasoning over video. CRN takes as input an…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Tanh Activation 설명 없음
Sigmoid Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

spatial-dccrn: dccrn equipped with frame-level angle feature and hybrid filtering for multi-channel speech enhancement

2022-10-17 · Shubo Lv, Yihui Fu, Yukai Jv, Lei Xie 외

Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network ba…

DenoisingSpeech Enhancement

Multi-channel end-to-end neural network for speech enhancement, source localization, and voice activity detection

2022-06-20 · Yuan Chen, Yicheng Hsu, Mingsian R. Bai

Speech enhancement and source localization has been active research for several decades with a wide range of real-world applications. Recently, the Deep Complex Convolution Recurrent network (DCCRN) has yielded impressiv…

Action DetectionActivity DetectionSpeech Enhancement

A Deep Representation Learning-based Speech Enhancement Method Using Complex Convolution Recurrent Variational Autoencoder

2023-12-15 · Yang Xiang, Jingguang Tian, Xinhui Hu, Xinkang Xu 외

Generally, the performance of deep neural networks (DNNs) heavily depends on the quality of data representation learning. Our preliminary work has emphasized the significance of deep representation learning (DRL) in the …

Representation LearningSpeech Enhancement

Causal Signal-Based DCCRN with Overlapped-Frame Prediction for Online Speech Enhancement

2023-09-07 · Julitta Bartolewska, Stanisław Kacprzak, Konrad Kowalczyk

The aim of speech enhancement is to improve speech signal quality and intelligibility from a noisy microphone signal. In many applications, it is crucial to enable processing with small computational complexity and minim…

Speech Enhancement

S-DCCRN: Super Wide Band DCCRN with learnable complex feature for speech enhancement

2021-11-16 · Shubo Lv, Yihui Fu, Mengtao Xing, Jiayao Sun 외

In speech enhancement, complex neural network has shown promising performance due to their effectiveness in processing complex-valued spectrum. Most of the recent speech enhancement approaches mainly focus on wide-band s…

16kDenoisingSpeech DenoisingSpeech Enhancement