Remarks on Optimal Scores for Speaker Recognition
In this article, we first establish the theory of optimal scores for speaker recognition. Our analysis shows that the minimum Bayes risk (MBR) decisions for both the speaker identification and speaker verification tasks can be based on a normalized likelihood (NL). When the underlying generative model is a linear Gaussian, the NL score is mathematically equivalent to the PLDA likelihood ratio, and the empirical scores based on cosine distance and Euclidean distance can be seen as approximations of this linear Gaussian NL score under some conditions. We discuss a number of properties of the NL score and perform a simple simulation experiment to demonstrate the properties of the NL score.
Code (0)
등록된 구현이 없습니다.
Tasks
Speaker IdentificationSpeaker RecognitionSpeaker VerificationSimilar Papers 제목 키워드 기반
Investigating Confidence Estimation Measures for Speaker Diarization
Speaker diarization systems segment a conversation recording based on the speakers' identity. Such systems can misclassify the speaker of a portion of audio due to a variety of factors, such as speech pattern variation, …
speaker-diarizationSpeaker Diarizationspeech-recognitionSpeech RecognitionDr-Vectors: Decision Residual Networks and an Improved Loss for Speaker Recognition
Many neural network speaker recognition systems model each speaker using a fixed-dimensional embedding vector. These embeddings are generally compared using either linear or 2nd-order scoring and, until recently, do not …
Speaker RecognitionLexical Bias In Essay Level Prediction
Automatically predicting the level of non-native English speakers given their written essays is an interesting machine learning problem. In this work I present the system "balikasg" that achieved the state-of-the-art per…
BIG-bench Machine LearningFeature EngineeringModel SelectionPredictionA Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding
Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR). However, they have not been totally proven to produce better performance on tasks othe…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Recognitionintent-classification+8The Phonexia VoxCeleb Speaker Recognition Challenge 2021 System Description
We describe the Phonexia submission for the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21) in the unsupervised speaker verification track. Our solution was very similar to IDLab's winning submission for VoxSRC-2…
ClusteringContrastive LearningSpeaker RecognitionSpeaker Verification