Orthonormal Embedding-based Deep Clustering for Single-channel Speech Separation
Deep clustering is a deep neural network-based speech separation algorithm that first trains the mixed component of signals with high-dimensional embeddings, and then uses a clustering algorithm to separate each mixture of sources. In this paper, we extend the baseline criterion of deep clustering with an additional regularization term to further improve the overall performance. This term plays a role in assigning a condition to the embeddings such that it gives less correlation to each embedding dimension, leading to better decomposition of the spectral bins. The regularization term helps to mitigate the unavoidable permutation problem in the conventional deep clustering method, which enables to bring better clustering through the formation of optimal embeddings. We evaluate the results by varying embedding dimension, signal-to-interference ratio (SIR), and gender dependency. The performance comparison with the source separation measurement metric, i.e. signal-to-distortion ratio (SDR), confirms that the proposed method outperforms the conventional deep clustering method.
Code (0)
등록된 구현이 없습니다.
Tasks
ClusteringDeep ClusteringSpeech SeparationSimilar Papers 제목 키워드 기반
Improved MVDR Beamforming Using LSTM Speech Models to Clean Spatial Clustering Masks
Spatial clustering techniques can achieve significant multi-channel noise reduction across relatively arbitrary microphone configurations, but have difficulty incorporating a detailed speech/noise model. In contrast, LST…
ClusteringCombining Spatial Clustering with LSTM Speech Models for Multichannel Speech Enhancement
Recurrent neural networks using the LSTM architecture can achieve significant single-channel noise reduction. It is not obvious, however, how to apply them to multi-channel inputs in a way that can generalize to new micr…
ClusteringSpeech EnhancementSpatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed system requires neither access to multi-c…
SegmentationEnhancement of Spatial Clustering-Based Time-Frequency Masks using LSTM Neural Networks
Recent works have shown that Deep Recurrent Neural Networks using the LSTM architecture can achieve strong single-channel speech enhancement by estimating time-frequency masks. However, these models do not naturally gene…
ClusteringSpeech EnhancementThe Speed Submission to DIHARD II: Contributions & Lessons Learned
This paper describes the speaker diarization systems developed for the Second DIHARD Speech Diarization Challenge (DIHARD II) by the Speed team. Besides describing the system, which considerably outperformed the challeng…
Action DetectionActivity DetectionClusteringspeaker-diarization+2