Single microphone speaker extraction using unified time-frequency Siamese-Unet
In this paper we present a unified time-frequency method for speaker extraction in clean and noisy conditions. Given a mixed signal, along with a reference signal, the common approaches for extracting the desired speaker are either applied in the time-domain or in the frequency-domain. In our approach, we propose a Siamese-Unet architecture that uses both representations. The Siamese encoders are applied in the frequency-domain to infer the embedding of the noisy and reference spectra, respectively. The concatenated representations are then fed into the decoder to estimate the real and imaginary components of the desired speaker, which are then inverse-transformed to the time-domain. The model is trained with the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) loss to exploit the time-domain information. The time-domain loss is also regularized with frequency-domain loss to preserve the speech patterns. Experimental results demonstrate that the unified approach is not only very easy to train, but also provides superior results as compared with state-of-the-art (SOTA) Blind Source Separation (BSS) methods, as well as commonly used speaker extraction approach.
Code (0)
등록된 구현이 없습니다.
Tasks
blind source separationDecoderSimilar Papers 제목 키워드 기반
Deep Ad-hoc Beamforming Based on Speaker Extraction for Target-Dependent Speech Separation
Recently, the research on ad-hoc microphone arrays with deep learning has drawn much attention, especially in speech enhancement and separation. Because an ad-hoc microphone array may cover such a large area that multipl…
channel selectionDeep LearningSpeech EnhancementSpeech SeparationBeamformer-Guided Target Speaker Extraction
We propose a Beamformer-guided Target Speaker Extraction (BG-TSE) method to extract a target speaker's voice from a multi-channel recording informed by the direction of arrival of the target. The proposed method employs …
Target Speaker ExtractionSpeaker Diarization and Identification from Single-Channel Classroom Audio Recording Using Virtual Microphones
Speaker identification in noisy audio recordings, specifically those from collaborative learning environments, can be extremely challenging. There is a need to identify individual students talking in small groups from ot…
speaker-diarizationSpeaker DiarizationSpeaker IdentificationSpeaker activity driven neural speech extraction
Target speech extraction, which extracts the speech of a target speaker in a mixture given auxiliary speaker clues, has recently received increased interest. Various clues have been investigated such as pre-recorded enro…
Speech ExtractionImproving speaker discrimination of target speech extraction with time-domain SpeakerBeam
Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation u…
Speaker IdentificationSpeech ExtractionSpeech Separation