Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?
Previous work on speaker adaptation for end-to-end speech synthesis still falls short in speaker similarity. We investigate an orthogonal approach to the current speaker adaptation paradigms, speaker augmentation, by creating artificial speakers and by taking advantage of low-quality data. The base Tacotron2 model is modified to account for the channel and dialect factors inherent in these corpora. In addition, we describe a warm-start training strategy that we adopted for Tacotron2 training. A large-scale listening test is conducted, and a distance metric is adopted to evaluate synthesis of dialects. This is followed by an analysis on synthesis quality, speaker and dialect similarity, and a remark on the effectiveness of our speaker augmentation approach. Audio samples are available online.
Code (1)
Tasks
Speech SynthesisSimilar Papers 제목 키워드 기반
Speaker verification-derived loss and data augmentation for DNN-based multispeaker speech synthesis
Building multispeaker neural network-based text-to-speech synthesis systems commonly relies on the availability of large amounts of high quality recordings from each speaker and conditioning the training process on the s…
Data AugmentationSpeaker VerificationSpeech Synthesistext-to-speech+2ASR data augmentation in low-resource settings using cross-lingual multi-speaker TTS and cross-lingual voice conversion
We explore cross-lingual multi-speaker speech synthesis and cross-lingual voice conversion applied to data augmentation for automatic speech recognition (ASR) systems in low/medium-resource scenarios. Through extensive e…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data Augmentationspeech-recognition+3Cross-speaker style transfer for text-to-speech using data augmentation
We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive data from a target speaker and supporting…
Data AugmentationStyle Transfertext-to-speechText to Speech+1Relational Data Selection for Data Augmentation of Speaker-dependent Multi-band MelGAN Vocoder
Nowadays, neural vocoders can generate very high-fidelity speech when a bunch of training data is available. Although a speaker-dependent (SD) vocoder usually outperforms a speaker-independent (SI) vocoder, it is impract…
Data AugmentationSpeaker VerificationUsing Data Augmentations and VTLN to Reduce Bias in Dutch End-to-End Speech Recognition Systems
Speech technology has improved greatly for norm speakers, i.e., adult native speakers of a language without speech impediments or strong accents. However, non-norm or diverse speaker groups show a distinct performance ga…
AnatomyData Augmentationspeech-recognitionSpeech Recognition