Meta-Learning Framework for End-to-End Imposter Identification in Unseen Speaker Recognition
Speaker identification systems are deployed in diverse environments, often different from the lab conditions on which they are trained and tested. In this paper, first, we show the problem of generalization using fixed thresholds (computed using EER metric) for imposter identification in unseen speaker recognition and then introduce a robust speaker-specific thresholding technique for better performance. Secondly, inspired by the recent use of meta-learning techniques in speaker verification, we propose an end-to-end meta-learning framework for imposter detection which decouples the problem of imposter detection from unseen speaker identification. Thus, unlike most prior works that use some heuristics to detect imposters, the proposed network learns to detect imposters by leveraging the utterances of the enrolled speakers. Furthermore, we show the efficacy of the proposed techniques on VoxCeleb1, VCTK and the FFSVC 2022 datasets, beating the baselines by up to 10%.
Code (0)
등록된 구현이 없습니다.
Tasks
Meta-LearningSpeaker IdentificationSpeaker RecognitionSpeaker VerificationSimilar Papers 제목 키워드 기반
Improved Relation Networks for End-to-End Speaker Verification and Identification
Speaker identification systems in a real-world scenario are tasked to identify a speaker amongst a set of enrolled speakers given just a few samples for each enrolled speaker. This paper demonstrates the effectiveness of…
Meta-LearningRelationSpeaker IdentificationSpeaker VerificationIncorporating Pass-Phrase Dependent Background Models for Text-Dependent Speaker Verification
In this paper, we propose pass-phrase dependent background models (PBMs) for text-dependent (TD) speaker verification (SV) to integrate the pass-phrase identification process into the conventional TD-SV system, where a P…
Speaker VerificationText-Dependent Speaker VerificationMeta-Learning for Short Utterance Speaker Recognition with Imbalance Length Pairs
In practical settings, a speaker recognition system needs to identify a speaker given a short utterance, while the enrollment utterance may be relatively long. However, existing speaker recognition models perform poorly …
Meta-LearningSpeaker IdentificationSpeaker RecognitionSpeaker VerificationMax-margin Metric Learning for Speaker Recognition
Probabilistic linear discriminant analysis (PLDA) is a popular normalization approach for the i-vector model, and has delivered state-of-the-art performance in speaker recognition. A potential problem of the PLDA model, …
Metric LearningSpeaker RecognitionComparison of Multiple Features and Modeling Methods for Text-dependent Speaker Verification
Text-dependent speaker verification is becoming popular in the speaker recognition society. However, the conventional i-vector framework which has been successful for speaker identification and other similar tasks works …
Speaker IdentificationSpeaker RecognitionSpeaker VerificationText-Dependent Speaker Verification+1