Short utterance compensation in speaker verification via cosine-based teacher-student learning of speaker embeddings
The short duration of an input utterance is one of the most critical threats that degrade the performance of speaker verification systems. This study aimed to develop an integrated text-independent speaker verification system that inputs utterances with short duration of 2 seconds or less. We propose an approach using a teacher-student learning framework for this goal, applied to short utterance compensation for the first time in our knowledge. The core concept of the proposed system is to conduct the compensation throughout the network that extracts the speaker embedding, mainly in phonetic-level, rather than compensating via a separate system after extracting the speaker embedding. In the proposed architecture, phonetic-level features where each feature represents a segment of 130 ms are extracted using convolutional layers. A layer of gated recurrent units extracts an utterance-level feature using phonetic-level features. The proposed approach also adopts a new objective function for teacher-student learning that considers both Kullback-Leibler divergence of output layers and cosine distance of speaker embeddings layers. Experiments were conducted using deep neural networks that take raw waveforms as input, and output speaker embeddings on VoxCeleb1 dataset. The proposed model could compensate approximately 65 \% of the performance degradation due to the shortened duration.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker VerificationText-Independent Speaker VerificationSimilar Papers 제목 키워드 기반
I-vector Transformation Using Conditional Generative Adversarial Networks for Short Utterance Speaker Verification
I-vector based text-independent speaker verification (SV) systems often have poor performance with short utterances, as the biased phonetic distribution in a short utterance makes the extracted i-vector unreliable. This …
Generative Adversarial NetworkSpeaker VerificationText-Independent Speaker VerificationAttention Back-end for Automatic Speaker Verification with Multiple Enrollment Utterances
Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of mult…
Speaker VerificationDeep Speaker: an End-to-End Neural Speaker Embedding System
We present Deep Speaker, a neural speaker embedding system that maps utterances to a hypersphere where speaker similarity is measured by cosine similarity. The embeddings generated by Deep Speaker can be used for many ta…
ClusteringSpeaker IdentificationSpeaker RecognitionTripletNovel Quality Metric for Duration Variability Compensation in Speaker Verification using i-Vectors
Automatic speaker verification (ASV) is the process to recognize persons using voice as biometric. The ASV systems show considerable recognition performance with sufficient amount of speech from matched condition. One of…
Speaker VerificationSegment Aggregation for short utterances speaker verification using raw waveforms
Most studies on speaker verification systems focus on long-duration utterances, which are composed of sufficient phonetic information. However, the performances of these systems are known to degrade when short-duration u…
Speaker Verification