Unsupervised Representation Learning for Speaker Recognition via Contrastive Equilibrium Learning
In this paper, we propose a simple but powerful unsupervised learning method for speaker recognition, namely Contrastive Equilibrium Learning (CEL), which increases the uncertainty on nuisance factors latent in the embeddings by employing the uniformity loss. Also, to preserve speaker discriminability, a contrastive similarity loss function is used together. Experimental results showed that the proposed CEL significantly outperforms the state-of-the-art unsupervised speaker verification systems and the best performing model achieved 8.01% and 4.01% EER on VoxCeleb1 and VOiCES evaluation sets, respectively. On top of that, the performance of the supervised speaker embedding networks trained with initial parameters pre-trained via CEL showed better performance than those trained with randomly initialized parameters.
Code (1)
Tasks
Representation LearningSpeaker RecognitionSpeaker VerificationSimilar Papers 제목 키워드 기반
Bootstrap Equilibrium and Probabilistic Speaker Representation Learning for Self-supervised Speaker Verification
In this paper, we propose self-supervised speaker representation learning strategies, which comprise of a bootstrap equilibrium speaker representation learning in the front-end and an uncertainty-aware probabilistic spea…
Contrastive LearningRepresentation LearningSpeaker VerificationMomentum Contrast Speaker Representation Learning
Unsupervised representation learning has shown remarkable achievement by reducing the performance gap with supervised feature learning, especially in the image domain. In this study, to extend the technique of unsupervis…
Contrastive LearningMetric LearningRepresentation LearningSpeaker Recognition+1Non-Contrastive Self-Supervised Learning of Utterance-Level Speech Representations
Considering the abundance of unlabeled speech data and the high labeling costs, unsupervised learning methods can be essential for better system development. One of the most successful methods is contrastive self-supervi…
Emotion RecognitionSelf-Supervised LearningSpeaker VerificationRobust Speaker Recognition with Transformers Using wav2vec 2.0
Recent advances in unsupervised speech representation learning discover new approaches and provide new state-of-the-art for diverse types of speech processing tasks. This paper presents an investigation of using wav2vec …
Data AugmentationRepresentation LearningSpeaker RecognitionSpeaker Verification+1Augmentation adversarial training for self-supervised speaker recognition
The goal of this work is to train robust speaker recognition models without speaker labels. Recent works on unsupervised speaker representations are based on contrastive learning in which they encourage within-utterance …
Contrastive LearningSpeaker Recognition