D²Net: A Denoising and Dereverberation Network Based on Two-branch Encoder and Dual-path Transformer
The simultaneous denoising and dereverberation for single-channel mixture speech under the complicated acoustic environment is considered to be a challengeable task. In this paper, we propose a denoising and dereverberation network named as D²Net in which a two-branch encoder (TBE) is designed to extract and selectively fuse features with different granularity. In addition, we design a global-local dual-path transformer (GLDPT) which introduces the local dense synthesizer attention (LDSA) in the dual-path transformer to improve the perception of local information. We evaluated our proposed D²Net and conducted ablation studies on the VoiceBank+DEMAND and WHAMR! datasets. Meanwhile, we chose three types of data in the WHAMR! dataset to verify the ability of the D²Net on the tasks of denoising-only, dereverberation-only, and simultaneous denoising and dereverberation, respectively. Experimental results show that our proposed model outperforms the comparative models, and all achieve better performance on the tasks of simultaneous denoising and dereverberation, dereverberation-only, and denoising-only, while keeping a small number of network parameters.
Code (0)
등록된 구현이 없습니다.
Tasks
DenoisingSpeech EnhancementSimilar Papers 제목 키워드 기반
Real-time Single-channel Dereverberation and Separation with Time-domainAudio Separation Network
We investigate the recently proposed Time-domain Audio Sep-aration Network (TasNet) in the task of real-time single-channel speech dereverberation. Unlike systems that take time-frequency representation of the au…
DenoisingSpeech DereverberationSpeech SeparationUformer: A Unet based dilated complex & real dual-path conformer network for simultaneous speech enhancement and dereverberation
Complex spectrum and magnitude are considered as two major features of speech enhancement and dereverberation. Traditional approaches always treat these two features separately, ignoring their underlying relationship. In…
DecoderSpeech EnhancementSimultaneous Denoising and Dereverberation Using Deep Embedding Features
Monaural speech dereverberation is a very challenging task because no spatial cues can be used. When the additive noises exist, this task becomes more challenging. In this paper, we propose a joint training method for si…
ClusteringDeep ClusteringDenoisingSpeech Denoising+2Speech Dereverberation with A Reverberation Time Shortening Target
This work proposes a new learning target based on reverberation time shortening (RTS) for speech dereverberation. The learning target for dereverberation is usually set as the direct-path speech or optionally with some e…
DenoisingSpeech DenoisingSpeech DereverberationSpeech Dereverberation with a Reverberation Time Shortening Target
This work proposes a new learning target based on reverberation time shortening (RTS) for speech dereverberation. The learning target for dereverberation is usually set as the direct-path speech or optionally with some e…
DenoisingSpeech DenoisingSpeech Dereverberation