Using Optimal Ratio Mask as Training Target for Supervised Speech Separation
Supervised speech separation uses supervised learning algorithms to learn a mapping from an input noisy signal to an output target. With the fast development of deep learning, supervised separation has become the most important direction in speech separation area in recent years. For the supervised algorithm, training target has a significant impact on the performance. Ideal ratio mask is a commonly used training target, which can improve the speech intelligibility and quality of the separated speech. However, it does not take into account the correlation between noise and clean speech. In this paper, we use the optimal ratio mask as the training target of the deep neural network (DNN) for speech separation. The experiments are carried out under various noise environments and signal to noise ratio (SNR) conditions. The results show that the optimal ratio mask outperforms other training targets in general.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech SeparationSimilar Papers 제목 키워드 기반
Point2Mask: Point-supervised Panoptic Segmentation via Optimal Transport
Weakly-supervised image segmentation has recently attracted increasing research attentions, aiming to avoid the expensive pixel-wise labeling. In this paper, we present an effective method, namely Point2Mask, to achieve …
Image SegmentationPanoptic SegmentationSemantic SegmentationSkeleton2vec: A Self-supervised Learning Framework with Contextualized Target Representations for Skeleton Sequence
Self-supervised pre-training paradigms have been extensively explored in the field of skeleton-based action recognition. In particular, methods based on masked prediction have pushed the performance of pre-training to a …
Action RecognitionPredictionRepresentation LearningSelf-Supervised Learning+1Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation
We propose a method to automatically generate a domain- and task-adaptive maskings of the given text for self-supervised pre-training, such that we can effectively adapt the language model to a particular target task (e.…
Language ModelingLanguage ModellingQuestion Answeringreinforcement-learning+4Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input
Masked Autoencoders is a simple yet powerful self-supervised learning method. However, it learns representations indirectly by reconstructing masked input patches. Several methods learn representations directly by predic…
Audio ClassificationAudio TaggingKeyword SpottingKeyword Spotting on Google Speech Commands+3Unsupervised training of neural mask-based beamforming
We present an unsupervised training approach for a neural network-based mask estimator in an acoustic beamforming application. The network is trained to maximize a likelihood criterion derived from a spatial mixture mode…
speech-recognitionSpeech Recognition