Target Speech Extraction: Independent Vector Extraction Guided by Supervised Speaker Identification
This manuscript proposes a novel robust procedure for the extraction of a speaker of interest (SOI) from a mixture of audio sources. The estimation of the SOI is performed via independent vector extraction (IVE). Since the blind IVE cannot distinguish the target source by itself, it is guided towards the SOI via frame-wise speaker identification based on deep learning. Still, an incorrect speaker can be extracted due to guidance failings, especially when processing challenging data. To identify such cases, we propose a criterion for non-intrusively assessing the estimated speaker. It utilizes the same model as the speaker identification, so no additional training is required. When incorrect extraction is detected, we propose a ``deflation'' step in which the incorrect source is subtracted from the mixture and, subsequently, another attempt to extract the SOI is performed. The process is repeated until successful extraction is achieved. The proposed procedure is experimentally tested on artificial and real-world datasets containing challenging phenomena: source movements, reverberation, transient noise, or microphone failures. The method is compared with state-of-the-art blind algorithms as well as with current fully supervised deep learning-based methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker IdentificationSpeech ExtractionSimilar Papers 제목 키워드 기반
Target Speech Extraction Based on Blind Source Separation and X-vector-based Speaker Selection Trained with Data Augmentation
Extracting the desired speech from a mixture is a meaningful and challenging task. The end-to-end DNN-based methods, though attractive, face the problem of generalization. In this paper, we explore a sequential approach …
blind source separationData AugmentationSpeaker RecognitionSpeech ExtractionSimilarity-and-Independence-Aware Beamformer with Iterative Casting and Boost Start for Target Source Extraction Using Reference
Target source extraction is significant for improving human speech intelligibility and the speech recognition performance of computers. This study describes a method for target source extraction, called the similarity-an…
Speech Enhancementspeech-recognitionSpeech RecognitionDeep Ad-hoc Beamforming Based on Speaker Extraction for Target-Dependent Speech Separation
Recently, the research on ad-hoc microphone arrays with deep learning has drawn much attention, especially in speech enhancement and separation. Because an ad-hoc microphone array may cover such a large area that multipl…
channel selectionDeep LearningSpeech EnhancementSpeech SeparationFew-shot learning of new sound classes for target sound extraction
Target sound extraction consists of extracting the sound of a target acoustic event (AE) class from a mixture of AE sounds. It can be realized using a neural network that extracts the target sound conditioned on a 1-hot …
Few-Shot LearningTarget Sound ExtractionOn the Use of Different Feature Extraction Methods for Linear and Non Linear kernels
The speech feature extraction has been a key focus in robust speech recognition research; it significantly affects the recognition performance. In this paper, we first study a set of different features extraction methods…
Robust Speech RecognitionSpeaker Identificationspeech-recognitionSpeech Recognition