paper-with-me

홈 › Papers

MooseNet: A Trainable Metric for Synthesized Speech with a PLDA Module

2023-01-17 · Ondřej Plátek, Ondřej Dušek

We present MooseNet, a trainable speech metric that predicts the listeners' Mean Opinion Score (MOS). We propose a novel approach where the Probabilistic Linear Discriminative Analysis (PLDA) generative model is used on top of an embedding obtained from a self-supervised learning (SSL) neural network (NN) model. We show that PLDA works well with a non-finetuned SSL model when trained only on 136 utterances (ca. one minute training time) and that PLDA consistently improves various neural MOS prediction models, even state-of-the-art models with task-specific fine-tuning. Our ablation study shows PLDA training superiority over SSL model fine-tuning in a low-resource scenario. We also improve SSL model fine-tuning using a convenient optimizer choice and additional contrastive and multi-task training objectives. The fine-tuned MooseNet NN with the PLDA module achieves the best results, surpassing the SSL baseline on the VoiceMOS Challenge data.

📄 PDF Abstract BibTeX arXiv:2301.07087

Code (1)

oplatek/moosenet-plda 공식 구현 pytorch

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

Tied Probabilistic Linear Discriminant Analysis for Speech Recognition

2014-11-04 · Liang Lu, Steve Renals

Acoustic models using probabilistic linear discriminant analysis (PLDA) capture the correlations within feature vectors using subspaces which do not vastly expand the model. This allows high dimensional and correlated fe…

speech-recognitionSpeech Recognition

Toroidal Probabilistic Spherical Discriminant Analysis

2022-10-27 · Anna Silnova, Niko Brümmer, Albert Swart, Lukáš Burget

In speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring back-ends are commonly used, namely cosine scoring and PLDA. We have recently proposed PSDA, an analog to PLDA t…

FormSpeaker Recognition

Probabilistic embeddings for speaker diarization

2020-04-06 · Anna Silnova, Niko Brümmer, Johan Rohdin, Themos Stafylakis 외

Speaker embeddings (x-vectors) extracted from very short segments of speech have recently been shown to give competitive performance in speaker diarization. We generalize this recipe by extracting from each speech segmen…

Clusteringspeaker-diarizationSpeaker Diarization

Unsupervised Adaptation of SPLDA

2015-11-20 · Jesús Villalba

State-of-the-art speaker recognition relays on models that need a large amount of training data. This models are successful in tasks like NIST SRE because there is sufficient data available. However, in real applications…

speaker-diarizationSpeaker DiarizationSpeaker Recognition

ChildAugment: Data Augmentation Methods for Zero-Resource Children's Speaker Verification

2024-02-23 · Vishwanath Pratap Singh, Md Sahidullah, Tomi Kinnunen

The accuracy of modern automatic speaker verification (ASV) systems, when trained exclusively on adult data, drops substantially when applied to children's speech. The scarcity of children's speech corpora hinders fine-t…

Data AugmentationSpeaker Verification