Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
Domain gap often degrades the performance of speaker verification (SV) systems when the statistical distributions of training data and real-world test speech are mismatched. Channel variation, a primary factor causing this gap, is less addressed than other issues (e.g., noise). Although various domain adaptation algorithms could be applied to handle this domain gap problem, most algorithms could not take the complex distribution structure in domain alignment with discriminative learning. In this paper, we propose a novel unsupervised domain adaptation method, i.e., Joint Partial Optimal Transport with Pseudo Label (JPOT-PL), to alleviate the channel mismatch problem. Leveraging the geometric-aware distance metric of optimal transport in distribution alignment, we further design a pseudo label-based discriminative learning where the pseudo label can be regarded as a new type of soft speaker label derived from the optimal coupling. With the JPOT-PL, we carry out experiments on the SV channel adaptation task with VoxCeleb as the basis corpus. Experiments show our method reduces EER by over 10% compared with several state-of-the-art channel adaptation algorithms.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain AdaptationPseudo LabelSpeaker VerificationUnsupervised Domain AdaptationSimilar Papers 제목 키워드 기반
Source -Free Domain Adaptation for Speaker Verification in Data-Scarce Languages and Noisy Channels
Domain adaptation is often hampered by exceedingly small target datasets and inaccessible source data. These conditions are prevalent in speech verification, where privacy policies and/or languages with scarce speech res…
Domain AdaptationSource-Free Domain AdaptationSpeaker VerificationInterpretable Dysarthric Speaker Adaptation based on Optimal-Transport
This work addresses the mismatch problem between the distribution of training data (source) and testing data (target), in the challenging context of dysarthric speech recognition. We focus on Speaker Adaptation (SA) in c…
Domain Adaptationspeech-recognitionSpeech RecognitionOptimal Transport-based Adaptation in Dysarthric Speech Tasks
In many real-world applications, the mismatch between distributions of training data (source) and test data (target) significantly degrades the performance of machine learning algorithms. In speech data, causes of this m…
speech-recognitionSpeech RecognitionDynamic Kernels and Channel Attention for Low Resource Speaker Verification
State-of-the-art speaker verification frameworks have typically focused on developing models with increasingly deeper (more layers) and wider (number of channels) models to improve their verification performance. Instead…
Speaker VerificationSpeech EnhancementDiscrete optimal transport is a strong audio adversarial attack
In this paper, we investigate discrete optimal transport (DOT) as a black-box attack against modern automatic speaker verification (ASV) and anti-spoofing countermeasure (CM) systems. Our attack operates as a post-proces…
Speaker VerificationAdversarial Attack