Learning Repeatable Speech Embeddings Using An Intra-class Correlation Regularizer
A good supervised embedding for a specific machine learning task is only sensitive to changes in the label of interest and is invariant to other confounding factors. We leverage the concept of repeatability from measurement theory to describe this property and propose to use the intra-class correlation coefficient (ICC) to evaluate the repeatability of embeddings. We then propose a novel regularizer, the ICC regularizer, as a complementary component for contrastive losses to guide deep neural networks to produce embeddings with higher repeatability. We use simulated data to explain why the ICC regularizer works better on minimizing the intra-class variance than the contrastive loss alone. We implement the ICC regularizer and apply it to three speech tasks: speaker verification, voice style conversion, and a clinical application for detecting dysphonic voice. The experimental results demonstrate that adding an ICC regularizer can improve the repeatability of learned embeddings compared to only using the contrastive loss; further, these embeddings lead to improved performance in these downstream tasks.
Code (1)
Tasks
Speaker VerificationSimilar Papers 제목 키워드 기반
Residual Information in Deep Speaker Embedding Architectures
Speaker embeddings represent a means to extract representative vectorial representations from a speech signal such that the representation pertains to the speaker identity alone. The embeddings are commonly used to class…
Learning Speech Emotion Representations in the Quaternion Domain
The modeling of human emotion expression in speech signals is an important, yet challenging task. The high resource demand of speech emotion recognition models, combined with the the general scarcity of emotion-labelled …
Emotion RecognitionSpeech Emotion RecognitionWe Need Variations in Speech Generation: Sub-center Modelling for Speaker Embeddings
Modeling the rich prosodic variations inherent in human speech is essential for generating natural-sounding speech. While speaker embeddings are commonly used as conditioning inputs in personalized speech generation, the…
Speaker RecognitionSpeech SynthesisVoice ConversionMitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different indi…
Speaker DiarizationA multicenter study on radiomic features from T2‐weighted images of a customized MR pelvic phantom setting the basis for robust radiomic models in clinics
Purpose To investigate the repeatability and reproducibility of radiomic features extracted from MR images and provide a workflow to identify robust features. Methods T2‐weighted images of a pelvic phantom were acqu…