Target Speech Extraction Based on Blind Source Separation and X-vector-based Speaker Selection Trained with Data Augmentation
Extracting the desired speech from a mixture is a meaningful and challenging task. The end-to-end DNN-based methods, though attractive, face the problem of generalization. In this paper, we explore a sequential approach for target speech extraction by combining blind source separation (BSS) with the x-vector based speaker recognition (SR) module. Two promising BSS methods based on source independence assumption, independent low-rank matrix analysis (ILRMA) and multi-channel variational autoencoder (MVAE), are utilized and compared. ILRMA employs nonnegative matrix factorization (NMF) to capture spectral structures of source signals and MVAE utilizes the strong modeling power of deep neural networks (DNN). However, the investigation of MVAE has been limited to the training with very few speakers and the speech signals of test speakers are usually included. We extend the training of MVAE using clean speech signals of 500 speakers to evaluate its generalization to unseen speakers. To improve the correct extraction rate, two data augmentation strategies are implemented to train the SR module. The performance of the proposed cascaded approach is investigated with test data constructed with real room impulse responses under varied environments.
Code (1)
Tasks
blind source separationData AugmentationSpeaker RecognitionSpeech ExtractionSimilar Papers 제목 키워드 기반
USEF-TSE: Universal Speaker Embedding Free Target Speaker Extraction
Target speaker extraction aims to separate the voice of a specific speaker from mixed speech. Traditionally, this process has relied on extracting a speaker embedding from a reference speech, in which a speaker recogniti…
Speaker RecognitionSpeech SeparationTarget Speaker ExtractionTF-MLPNet: Tiny Real-Time Neural Speech Separation
Speech separation on hearable devices can enable transformative augmented and enhanced hearing capabilities. However, state-of-the-art speech separation networks cannot run in real-time on tiny, low-power neural accelera…
Speech ExtractionSpeech SeparationAnalysis of impact of emotions on target speech extraction and speech separation
Recently, the performance of blind speech separation (BSS) and target speech extraction (TSE) has greatly progressed. Most works, however, focus on relatively well-controlled conditions using, e.g., read speech. The perf…
Speaker VerificationSpeech ExtractionSpeech SeparationSimilarity-and-Independence-Aware Beamformer: Method for Target Source Extraction using Magnitude Spectrogram as Reference
This study presents a novel method for source extraction, referred to as the similarity-and-independence-aware beamformer (SIBF). The SIBF extracts the target signal using a rough magnitude spectrogram as the reference s…
Speech EnhancementIndependent Vector Extraction for Fast Joint Blind Source Separation and Dereverberation
We address a blind source separation (BSS) problem in a noisy reverberant environment in which the number of microphones $M$ is greater than the number of sources of interest, and the other noise components can be approx…
blind source separation