Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-supervised Speaker Verification
In this paper, we propose self-supervised speaker representation learning strategies, which comprise of a bootstrap equilibrium speaker representation learning in the front-end and an uncertainty-aware probabilistic speaker embedding training in the back-end. In the front-end stage, we learn the speaker representations via the bootstrap training scheme with the uniformity regularization term. In the back-end stage, the probabilistic speaker embeddings are estimated by maximizing the mutual likelihood score between the speech samples belonging to the same speaker, which provide not only speaker representations but also data uncertainty. Experimental results show that the proposed bootstrap equilibrium training strategy can effectively help learn the speaker representations and outperforms the conventional methods based on contrastive learning. Also, we demonstrate that the integrated two-stage framework further improves the speaker verification performance on the VoxCeleb1 test set in terms of EER and MinDCF.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive LearningRepresentation LearningSpeaker VerificationSimilar Papers 제목 키워드 기반
The DKU-DukeECE System for the Self-Supervision Speaker Verification Task of the 2021 VoxCeleb Speaker Recognition Challenge
This report describes the submission of the DKU-DukeECE team to the self-supervision speaker verification task of the 2021 VoxCeleb Speaker Recognition Challenge (VoxSRC). Our method employs an iterative labeling framewo…
Pseudo LabelRepresentation LearningSpeaker RecognitionSpeaker VerificationSelf-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling
Recent developments in Self-Supervised Learning (SSL) have demonstrated significant potential for Speaker Verification (SV), but closing the performance gap with supervised systems remains an ongoing challenge. Standard …
Data AugmentationSelf-Supervised LearningSpeaker VerificationSpeaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
Iterative self-training, or iterative pseudo-labeling (IPL) -- using an improved model from the current iteration to provide pseudo-labels for the next iteration -- has proven to be a powerful approach to enhance the qua…
Speaker RecognitionSpeaker VerificationSelf-supervised learning for robust voice cloning
Voice cloning is a difficult task which requires robust and informative features incorporated in a high quality TTS system in order to effectively copy an unseen speaker's voice. In our work, we utilize features learned …
Self-Supervised LearningSpeech SynthesisVoice CloningSelf-supervised Speaker Diarization
Over the last few years, deep learning has grown in popularity for speaker verification, identification, and diarization. Inarguably, a significant part of this success is due to the demonstrated effectiveness of their s…
speaker-diarizationSpeaker DiarizationSpeaker Verification