MooseNet: A Trainable Metric for Synthesized Speech with a PLDA Module
We present MooseNet, a trainable speech metric that predicts the listeners' Mean Opinion Score (MOS). We propose a novel approach where the Probabilistic Linear Discriminative Analysis (PLDA) generative model is used on top of an embedding obtained from a self-supervised learning (SSL) neural network (NN) model. We show that PLDA works well with a non-finetuned SSL model when trained only on 136 utterances (ca. one minute training time) and that PLDA consistently improves various neural MOS prediction models, even state-of-the-art models with task-specific fine-tuning. Our ablation study shows PLDA training superiority over SSL model fine-tuning in a low-resource scenario. We also improve SSL model fine-tuning using a convenient optimizer choice and additional contrastive and multi-task training objectives. The fine-tuned MooseNet NN with the PLDA module achieves the best results, surpassing the SSL baseline on the VoiceMOS Challenge data.
Code (1)
Tasks
Self-Supervised LearningSimilar Papers 제목 키워드 기반
Tied Probabilistic Linear Discriminant Analysis for Speech Recognition
Acoustic models using probabilistic linear discriminant analysis (PLDA) capture the correlations within feature vectors using subspaces which do not vastly expand the model. This allows high dimensional and correlated fe…
speech-recognitionSpeech RecognitionToroidal Probabilistic Spherical Discriminant Analysis
In speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring back-ends are commonly used, namely cosine scoring and PLDA. We have recently proposed PSDA, an analog to PLDA t…
FormSpeaker RecognitionProbabilistic embeddings for speaker diarization
Speaker embeddings (x-vectors) extracted from very short segments of speech have recently been shown to give competitive performance in speaker diarization. We generalize this recipe by extracting from each speech segmen…
Clusteringspeaker-diarizationSpeaker DiarizationUnsupervised Adaptation of SPLDA
State-of-the-art speaker recognition relays on models that need a large amount of training data. This models are successful in tasks like NIST SRE because there is sufficient data available. However, in real applications…
speaker-diarizationSpeaker DiarizationSpeaker RecognitionChildAugment: Data Augmentation Methods for Zero-Resource Children's Speaker Verification
The accuracy of modern automatic speaker verification (ASV) systems, when trained exclusively on adult data, drops substantially when applied to children's speech. The scarcity of children's speech corpora hinders fine-t…
Data AugmentationSpeaker Verification