Adaptive blind audio source extraction supervised by dominant speaker identification using x-vectors
We propose a novel algorithm for adaptive blind audio source extraction. The proposed method is based on independent vector analysis and utilizes the auxiliary function optimization to achieve high convergence speed. The algorithm is partially supervised by a pilot signal related to the source of interest (SOI), which ensures that the method correctly extracts the utterance of the desired speaker. The pilot is based on the identification of a dominant speaker in the mixture using x-vectors. The properties of the x-vectors computed in the presence of cross-talk are experimentally analyzed. The proposed approach is verified in a scenario with a moving SOI, static interfering speaker, and environmental noise.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker IdentificationSimilar Papers 제목 키워드 기반
Target Speech Extraction: Independent Vector Extraction Guided by Supervised Speaker Identification
This manuscript proposes a novel robust procedure for the extraction of a speaker of interest (SOI) from a mixture of audio sources. The estimation of the SOI is performed via independent vector extraction (IVE). Since t…
Speaker IdentificationSpeech ExtractionNeural Spectral Band Generation for Audio Coding
Audio bandwidth extension is the task of reconstructing missing high frequency components of bandwidth-limited audio signals, where bandwidth limitation is a common issue for audio signals due to several reasons, includi…
Bandwidth ExtensionGeneralized Canonical Correlation Analysis and Its Application to Blind Source Separation Based on a Dual-Linear Predictor Structure
Blind source separation (BSS) is one of the most important and established research topics in signal processing and many algorithms have been proposed based on different statistical properties of the source signals. For …
blind source separationArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior
Blind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures recorded by a microphone array. The problem is challenging because it is a blind inverse problem, i.e., the microphone array geom…
Room Impulse Response (RIR)Speech SeparationAudioSlots: A slot-centric generative model for audio separation
In a range of recent works, object-centric architectures have been shown to be suitable for unsupervised scene decomposition in the vision domain. Inspired by these methods we present AudioSlots, a slot-centric generativ…
blind source separationDecoderSpeech Separation