Deep Interaction between Masking and Mapping Targets for Single-Channel Speech Enhancement
The most recent deep neural network (DNN) models exhibit impressive denoising performance in the time-frequency (T-F) magnitude domain. However, the phase is also a critical component of the speech signal that is easily overlooked. In this paper, we propose a multi-branch dilated convolutional network (DCN) to simultaneously enhance the magnitude and phase of noisy speech. A causal and robust monaural speech enhancement system is achieved based on the multi-objective learning framework of the complex spectrum and the ideal ratio mask (IRM) targets. In the process of joint learning, the intermediate estimation of IRM targets is used as a way of generating feature attention factors to realize the information interaction between the two targets. Moreover, the proposed multi-scale dilated convolution enables the DCN model to have a more efficient temporal modeling capability. Experimental results show that compared with other state-of-the-art models, this model achieves better speech quality and intelligibility with less computation.
Code (0)
등록된 구현이 없습니다.
Tasks
DenoisingSpeech EnhancementMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions
In multimodal representation learning, synergistic interactions between modalities not only provide complementary information but also create unique outcomes through specific interaction patterns that no single modality …
Representation LearningInformation ExtractionMasked adversarial neural network for cell type deconvolution in spatial transcriptomics
Accurately determining cell type composition in disease-relevant tissues is crucial for identifying disease targets. Most existing spatial transcriptomics (ST) technologies cannot achieve single-cell resolution, making i…
The Dynamic Duo of Collaborative Masking and Target for Advanced Masked Autoencoder Learning
Masked autoencoders (MAE) have recently succeeded in self-supervised vision representation learning. Previous work mainly applied custom-designed (e.g., random, block-wise) masking or teacher (e.g., CLIP)-guided masking …
DecoderRepresentation LearningCross-Drone Transformer Network for Robust Single Object Tracking
Drones have been widely used in a variety of applications, e.g., aerial photography and military security, because of their high maneuverability and broad views compared with fixed cameras. Multi-drone tracking systems c…
ObjectObject TrackingVisual Object TrackingVisual TrackingNP-TCMtarget: a network pharmacology platform for exploring mechanisms of action of Traditional Chinese medicine
The biological targets of traditional Chinese medicine (TCM) are the core effectors mediating the interaction between TCM and the human body. Identification of TCM targets is essential to elucidate the chemical basis and…
Binary Classification