Powerful Speaker Embedding Training Framework by Adversarially Disentangled Identity Representation
The main challenge of speaker verification in the wild is the interference caused by irrelevant information in speech and the lack of speaker labels in speech datasets. In order to solve the above problems, we propose a novel speaker embedding training framework based on adversarially disentangled identity representation. Our key insight is to adversarially learn the identity-purified features for speaker verification, and learn an identity-irrelated feature whose speaker information cannot be distinguished. Based on the existing state-of-the-art speaker verification models, we improve them without adjusting the structure and hyper-parameters of any model. Experiments prove that the framework we propose can significantly improve the performance of speaker verification from the original model without any empirical adjustments. Proving that it is particularly useful for alleviating the lack of speaker labels.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker VerificationSimilar Papers 제목 키워드 기반
Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System
Speaker embedding based zero-shot Text-to-Speech (TTS) systems enable high-quality speech synthesis for unseen speakers using minimal data. However, these systems are vulnerable to adversarial attacks, where an attacker …
Adversarial PurificationSpeech Synthesistext-to-speechText to SpeechChannel adversarial training for speaker verification and diarization
Previous work has encouraged domain-invariance in deep speaker embedding by adversarially classifying the dataset or labelled environment to which the generated features belong. We propose a training strategy which aims …
Speaker VerificationDeep Representation Decomposition for Rate-Invariant Speaker Verification
While promising performance for speaker verification has been achieved by deep speaker embeddings, the advantage would reduce in the case of speaking-style variability. Speaking rate mismatch is often observed in practic…
Speaker VerificationMulti-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations
Recently, cycle-consistent adversarial network (Cycle-GAN) has been successfully applied to voice conversion to a different speaker without parallel data, although in those approaches an individual model is needed for ea…
DecoderVoice ConversionECAPA-TDNN Embeddings for Speaker Diarization
Learning robust speaker embeddings is a crucial step in speaker diarization. Deep neural networks can accurately capture speaker discriminative characteristics and popular deep embeddings such as x-vectors are nowadays a…
speaker-diarizationSpeaker DiarizationSpeaker Verification