Fast variational Bayes for heavy-tailed PLDA applied to i-vectors and x-vectors
The standard state-of-the-art backend for text-independent speaker recognizers that use i-vectors or x-vectors, is Gaussian PLDA (G-PLDA), assisted by a Gaussianization step involving length normalization. G-PLDA can be trained with both generative or discriminative methods. It has long been known that heavy-tailed PLDA (HT-PLDA), applied without length normalization, gives similar accuracy, but at considerable extra computational cost. We have recently introduced a fast scoring algorithm for a discriminatively trained HT-PLDA backend. This paper extends that work by introducing a fast, variational Bayes, generative training algorithm. We compare old and new backends, with and without length-normalization, with i-vectors and x-vectors, on SRE'10, SRE'16 and SITW.
Code (1)
Similar Papers 제목 키워드 기반
Unsupervised Adaptation of SPLDA
State-of-the-art speaker recognition relays on models that need a large amount of training data. This models are successful in tasks like NIST SRE because there is sufficient data available. However, in real applications…
speaker-diarizationSpeaker DiarizationSpeaker RecognitionBayesian SPLDA
In this document we are going to derive the equations needed to implement a Variational Bayes estimation of the parameters of the simplified probabilistic linear discriminant analysis (SPLDA) model. This can be used to a…
Gaussian meta-embeddings for efficient scoring of a heavy-tailed PLDA model
Embeddings in machine learning are low-dimensional representations of complex input patterns, with the property that simple geometric operations like Euclidean distances and dot products can be used for classification an…
Speaker RecognitionPosterior and variational inference for deep neural networks with heavy-tailed weights
We consider deep neural networks in a Bayesian framework with a prior distribution sampling the network weights at random. Following a recent idea of Agapiou and Castillo (2023), who show that heavy-tailed prior distribu…
Model SelectionVariational InferenceGeneralized domain adaptation framework for parametric back-end in speaker recognition
State-of-the-art speaker recognition systems comprise a speaker embedding front-end followed by a probabilistic linear discriminant analysis (PLDA) back-end. The effectiveness of these components relies on the availabili…
Domain AdaptationSpeaker RecognitionUnsupervised Domain Adaptation