paper-with-me

홈 › Papers

Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-supervised Speaker Verification

2021-12-16 · Sung Hwan Mun, Min Hyun Han, Dongjune Lee, JiHwan Kim, Nam Soo Kim

In this paper, we propose self-supervised speaker representation learning strategies, which comprise of a bootstrap equilibrium speaker representation learning in the front-end and an uncertainty-aware probabilistic speaker embedding training in the back-end. In the front-end stage, we learn the speaker representations via the bootstrap training scheme with the uniformity regularization term. In the back-end stage, the probabilistic speaker embeddings are estimated by maximizing the mutual likelihood score between the speech samples belonging to the same speaker, which provide not only speaker representations but also data uncertainty. Experimental results show that the proposed bootstrap equilibrium training strategy can effectively help learn the speaker representations and outperforms the conventional methods based on contrastive learning. Also, we demonstrate that the integrated two-stage framework further improves the speaker verification performance on the VoxCeleb1 test set in terms of EER and MinDCF.

📄 PDF Abstract BibTeX arXiv:2112.08929

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningRepresentation LearningSpeaker Verification

Similar Papers 제목 키워드 기반

The DKU-DukeECE System for the Self-Supervision Speaker Verification Task of the 2021 VoxCeleb Speaker Recognition Challenge

2021-09-07 · Danwei Cai, Ming Li

This report describes the submission of the DKU-DukeECE team to the self-supervision speaker verification task of the 2021 VoxCeleb Speaker Recognition Challenge (VoxSRC). Our method employs an iterative labeling framewo…

Pseudo LabelRepresentation LearningSpeaker RecognitionSpeaker Verification

Self-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling

2025-01-29 · Theo Lepage, Reda Dehak

Recent developments in Self-Supervised Learning (SSL) have demonstrated significant potential for Speaker Verification (SV), but closing the performance gap with supervised systems remains an ongoing challenge. Standard …

Data AugmentationSelf-Supervised LearningSpeaker Verification

Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels

2024-09-16 · Zakaria Aldeneh, Takuya Higuchi, Jee-weon Jung, Li-Wei Chen 외

Iterative self-training, or iterative pseudo-labeling (IPL) -- using an improved model from the current iteration to provide pseudo-labels for the next iteration -- has proven to be a powerful approach to enhance the qua…

Speaker RecognitionSpeaker Verification

Self-supervised learning for robust voice cloning

2022-04-07 · Konstantinos Klapsas, Nikolaos Ellinas, Karolos Nikitaras, Georgios Vamvoukakis 외

Voice cloning is a difficult task which requires robust and informative features incorporated in a high quality TTS system in order to effectively copy an unseen speaker's voice. In our work, we utilize features learned …

Self-Supervised LearningSpeech SynthesisVoice Cloning

Self-supervised Speaker Diarization

2022-04-08 · Yehoshua Dissen, Felix Kreuk, Joseph Keshet

Over the last few years, deep learning has grown in popularity for speaker verification, identification, and diarization. Inarguably, a significant part of this success is due to the demonstrated effectiveness of their s…

speaker-diarizationSpeaker DiarizationSpeaker Verification