Multi-Label Training for Text-Independent Speaker Identification
In this paper, we propose a novel strategy for text-independent speaker identification system: Multi-Label Training (MLT). Instead of the commonly used one-to-one correspondence between the speech and the speaker label, we divide all the speeches of each speaker into several subgroups, with each subgroup assigned a different set of labels. During the identification process, a specific speaker is identified as long as the predicted label is the same as one of his/her corresponding labels. We found that this method can force the model to distinguish the data more accurately, and somehow takes advantages of ensemble learning, while avoiding the significant increase of computation and storage burden. In the experiments, we found that not only in clean conditions, but also in noisy conditions with speech enhancement, Multi-Label Training can still achieve better identification performance than commom methods. It should be noted that the proposed strategy can be easily applied to almost all current text-independent speaker identification models to achieve further improvements.
Code (0)
등록된 구현이 없습니다.
Tasks
Ensemble LearningSpeaker IdentificationSpeech EnhancementSimilar Papers 제목 키워드 기반
Controllable speech synthesis by learning discrete phoneme-level prosodic representations
In this paper, we present a novel method for phoneme-level prosody control of F0 and duration using intuitive discrete labels. We propose an unsupervised prosodic clustering process which is used to discretize phoneme-le…
ClusteringSpeech Synthesistext-to-speechText to SpeechSpeakerStew: Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification System
In this paper, we describe SpeakerStew - a hybrid system to perform speaker verification on 46 languages. Two core ideas were explored in this system: (1) Pooling training data of different languages together for multili…
Speaker RecognitionSpeaker VerificationText-Independent Speaker VerificationJukeBox: A Multilingual Singer Recognition Dataset
A text-independent speaker recognition system relies on successfully encoding speech factors such as vocal pitch, intensity, and timbre to achieve good performance. A majority of such systems are trained and evaluated us…
Speaker RecognitionText-Independent Speaker RecognitionDisentangling Correlated Speaker and Noise for Speech Synthesis via Data Augmentation and Adversarial Factorization
To leverage crowd-sourced data to train multi-speaker text-to-speech (TTS) models that can synthesize clean speech for all speakers, it is essential to learn disentangled representations which can independently control t…
Data AugmentationDisentanglementSpeech Synthesistext-to-speech+1Effect of different splitting criteria on the performance of speech emotion recognition
Traditional speech emotion recognition (SER) evaluations have been performed merely on a speaker-independent condition; some of them even did not evaluate their result on this condition. This paper highlights the importa…
Emotion RecognitionSentenceSpeech Emotion Recognition