VoiceExtender: Short-utterance Text-independent Speaker Verification with Guided Diffusion Model
Speaker verification (SV) performance deteriorates as utterances become shorter. To this end, we propose a new architecture called VoiceExtender which provides a promising solution for improving SV performance when handling short-duration speech signals. We use two guided diffusion models, the built-in and the external speaker embedding (SE) guided diffusion model, both of which utilize a diffusion model-based sample generator that leverages SE guidance to augment the speech features based on a short utterance. Extensive experimental results on the VoxCeleb1 dataset show that our method outperforms the baseline, with relative improvements in equal error rate (EER) of 46.1%, 35.7%, 10.4%, and 5.7% for the short utterance conditions of 0.5, 1.0, 1.5, and 2.0 seconds, respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker VerificationText-Independent Speaker VerificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
MFA: TDNN with Multi-scale Frequency-channel Attention for Text-independent Speaker Verification with Short Utterances
The time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteris…
Speaker VerificationText-Independent Speaker VerificationI-vector Transformation Using Conditional Generative Adversarial Networks for Short Utterance Speaker Verification
I-vector based text-independent speaker verification (SV) systems often have poor performance with short utterances, as the biased phonetic distribution in a short utterance makes the extracted i-vector unreliable. This …
Generative Adversarial NetworkSpeaker VerificationText-Independent Speaker VerificationShort utterance compensation in speaker verification via cosine-based teacher-student learning of speaker embeddings
The short duration of an input utterance is one of the most critical threats that degrade the performance of speaker verification systems. This study aimed to develop an integrated text-independent speaker verification s…
Speaker VerificationText-Independent Speaker VerificationAn Empirical Study on Text-Independent Speaker Verification based on the GE2E Method
While many researchers in the speaker recognition area have started to replace the former classical state-of-the-art methods with deep learning techniques, some of the traditional i-vector-based methods are still state-o…
Speaker RecognitionSpeaker VerificationText-Independent Speaker VerificationDeep neural network based i-vector mapping for speaker verification using short utterances
Text-independent speaker recognition using short utterances is a highly challenging task due to the large variation and content mismatch between short utterances. I-vector based systems have become the standard in speake…
Speaker RecognitionSpeaker VerificationText-Independent Speaker Recognition