Centroid-based deep metric learning for speaker recognition
Speaker embedding models that utilize neural networks to map utterances to a space where distances reflect similarity between speakers have driven recent progress in the speaker recognition task. However, there is still a significant performance gap between recognizing speakers in the training set and unseen speakers. The latter case corresponds to the few-shot learning task, where a trained model is evaluated on unseen classes. Here, we optimize a speaker embedding model with prototypical network loss (PNL), a state-of-the-art approach for the few-shot image classification task. The resulting embedding model outperforms the state-of-the-art triplet loss based models in both speaker verification and identification tasks, for both seen and unseen speakers.
Code (0)
등록된 구현이 없습니다.
Tasks
Few-Shot Image ClassificationFew-Shot LearningGeneral Classificationimage-classificationImage ClassificationMetric LearningSpeaker RecognitionSpeaker VerificationTripletMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Frequency-centroid features for word recognition of non-native English speakers
The objective of this work is to investigate complementary features which can aid the quintessential Mel frequency cepstral coefficients (MFCCs) in the task of closed, limited set word recognition for non-native English …
Audio-to-Image Encoding for Improved Voice Characteristic Detection Using Deep Convolutional Neural Networks
This paper introduces a novel audio-to-image encoding framework that integrates multiple dimensions of voice characteristics into a single RGB image for speaker recognition. In this method, the green channel encodes raw …
Speaker RecognitionAbsolute decision corrupts absolutely: conservative online speaker diarisation
Our focus lies in developing an online speaker diarisation framework which demonstrates robust performance across diverse domains. In online speaker diarisation, outputs generated in real-time are irreversible, and a few…
Visual Recognition with Deep Nearest Centroids
We devise deep nearest centroids (DNC), a conceptually elegant yet surprisingly effective network for large-scale visual recognition, by revisiting Nearest Centroids, one of the most classic and simple classifiers. Curre…
Decision Makingimage-classificationImage ClassificationObject RecognitionLength- and Noise-aware Training Techniques for Short-utterance Speaker Recognition
Speaker recognition performance has been greatly improved with the emergence of deep learning. Deep neural networks show the capacity to effectively deal with impacts of noise and reverberation, making them attractive to…
Representation LearningSpeaker Recognition