Speaker verification-derived loss and data augmentation for DNN-based multispeaker speech synthesis
Building multispeaker neural network-based text-to-speech synthesis systems commonly relies on the availability of large amounts of high quality recordings from each speaker and conditioning the training process on the speaker's identity or on a learned representation of it. However, when little data is available from each speaker, or the number of speakers is limited, the multispeaker TTS can be hard to train and will result in poor speaker similarity and naturalness. In order to address this issue, we explore two directions: forcing the network to learn a better speaker identity representation by appending an additional loss term; and augmenting the input data pertaining to each speaker using waveform manipulation methods. We show that both methods are efficient when evaluated with both objective and subjective measures. The additional loss term aids the speaker similarity, while the data augmentation improves the intelligibility of the multispeaker TTS system.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationSpeaker VerificationSpeech Synthesistext-to-speechText to SpeechText-To-Speech SynthesisSimilar Papers 제목 키워드 기반
DASA: Difficulty-Aware Semantic Augmentation for Speaker Verification
Data augmentation is vital to the generalization ability and robustness of deep neural networks (DNNs) models. Existing augmentation methods for speaker verification manipulate the raw signal, which are time-consuming an…
Data AugmentationDiversitySpeaker VerificationAsymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification
Contrastive self-supervised learning (CSL) for speaker verification (SV) has drawn increasing interest recently due to its ability to exploit unlabeled data. Performing data augmentation on raw waveforms, such as adding …
Data AugmentationSelf-Supervised LearningSpeaker VerificationGetting More for Less: Using Weak Labels and AV-Mixup for Robust Audio-Visual Speaker Verification
Distance Metric Learning (DML) has typically dominated the audio-visual speaker verification problem space, owing to strong performance in new and unseen classes. In our work, we explored multitask learning techniques to…
Metric LearningMulti-Task LearningSpeaker VerificationContrastive-mixup learning for improved speaker verification
This paper proposes a novel formulation of prototypical loss with mixup for speaker verification. Mixup is a simple yet efficient data augmentation technique that fabricates a weighted combination of random data point an…
Data AugmentationMetric LearningSpeaker VerificationA Comparison of Metric Learning Loss Functions for End-To-End Speaker Verification
Despite the growing popularity of metric learning approaches, very little work has attempted to perform a fair comparison of these techniques for speaker verification. We try to fill this gap and compare several metric l…
Metric LearningSpeaker VerificationTriplet