Bayesian Learning for Domain-Invariant Speaker Verification and Anti-Spoofing
The performance of automatic speaker verification (ASV) and anti-spoofing drops seriously under real-world domain mismatch conditions. The relaxed instance frequency-wise normalization (RFN), which normalizes the frequency components based on the feature statistics along the time and channel axes, is a promising approach to reducing the domain dependence in the feature maps of a speaker embedding network. We advocate that the different frequencies should receive different weights and that the weights' uncertainty due to domain shift should be accounted for. To these ends, we propose leveraging variational inference to model the posterior distribution of the weights, which results in Bayesian weighted RFN (BWRFN). This approach overcomes the limitations of fixed-weight RFN, making it more effective under domain mismatch conditions. Extensive experiments on cross-dataset ASV, cross-TTS anti-spoofing, and spoofing-robust ASV show that BWRFN is significantly better than WRFN and RFN.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker VerificationVariational InferenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Bayesian x-vector: Bayesian Neural Network based x-vector System for Speaker Verification
Speaker verification systems usually suffer from the mismatch problem between training and evaluation data, such as speaker population mismatch, the channel and environment variations. In order to address this issue, it …
Speaker VerificationGenerative Adversarial Speaker Embedding Networks for Domain Robust End-to-End Speaker Verification
This article presents a novel approach for learning domain-invariant speaker embeddings using Generative Adversarial Networks. The main idea is to confuse a domain discriminator so that is can't tell if embeddings are fr…
Dimensionality ReductionSpeaker VerificationNoise Invariant Frame Selection: A Simple Method to Address the Background Noise Problem for Text-independent Speaker Verification
The performance of speaker-related systems usually degrades heavily in practical applications largely due to the presence of background noise. To improve the robustness of such systems in unknown noisy environments, this…
QuantizationSpeaker VerificationText-Independent Speaker VerificationPrototype and Instance Contrastive Learning for Unsupervised Domain Adaptation in Speaker Verification
Speaker verification system trained on one domain usually suffers performance degradation when applied to another domain. To address this challenge, researchers commonly use feature distribution matching-based methods in…
Contrastive LearningDomain AdaptationSpeaker VerificationUnsupervised Domain AdaptationChannel adversarial training for speaker verification and diarization
Previous work has encouraged domain-invariance in deep speaker embedding by adversarially classifying the dataset or labelled environment to which the generated features belong. We propose a training strategy which aims …
Speaker Verification