Feature Enhancement with Deep Feature Losses for Speaker Verification
Speaker Verification still suffers from the challenge of generalization to novel adverse environments. We leverage on the recent advancements made by deep learning based speech enhancement and propose a feature-domain supervised denoising based solution. We propose to use Deep Feature Loss which optimizes the enhancement network in the hidden activation space of a pre-trained auxiliary speaker embedding network. We experimentally verify the approach on simulated and real data. A simulated testing setup is created using various noise types at different SNR levels. For evaluation on real data, we choose BabyTrain corpus which consists of children recordings in uncontrolled environments. We observe consistent gains in every condition over the state-of-the-art augmented Factorized-TDNN x-vector system. On BabyTrain corpus, we observe relative gains of 10.38% and 12.40% in minDCF and EER respectively.
Code (1)
Tasks
DenoisingSpeaker VerificationSpeech EnhancementSimilar Papers 제목 키워드 기반
Extended U-Net for Speaker Verification in Noisy Environments
Background noise is a well-known factor that deteriorates the accuracy and reliability of speaker verification (SV) systems by blurring speech intelligibility. Various studies have used separate pretrained enhancement mo…
DenoisingSpeaker IdentificationSpeaker VerificationVoiceID Loss: Speech Enhancement for Speaker Verification
In this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly used loss functions for speech enhancem…
Speaker VerificationSpeech EnhancementDeep multi-metric learning for text-independent speaker verification
Text-independent speaker verification is an important artificial intelligence problem that has a wide spectrum of applications, such as criminal investigation, payment certification, and interest-based customer services.…
Metric LearningSpeaker VerificationText-Independent Speaker VerificationTripletMulti-Channel Speaker Verification for Single and Multi-talker Speech
To improve speaker verification in real scenarios with interference speakers, noise, and reverberation, we propose to bring together advancements made in multi-channel speech features. Specifically, we combine spectral, …
Action DetectionActivity DetectionSpeaker VerificationSpeech Enhancement+1Towards Robust Speaker Verification with Target Speaker Enhancement
This paper proposes the target speaker enhancement based speaker verification network (TASE-SVNet), an all neural model that couples target speaker enhancement and speaker embedding extraction for robust speaker verifica…
Speaker VerificationSpeech Enhancement