Latent Filling: Latent Space Data Augmentation for Zero-shot Speech Synthesis
Previous works in zero-shot text-to-speech (ZS-TTS) have attempted to enhance its systems by enlarging the training data through crowd-sourcing or augmenting existing speech data. However, the use of low-quality data has led to a decline in the overall system performance. To avoid such degradation, instead of directly augmenting the input data, we propose a latent filling (LF) method that adopts simple but effective latent space data augmentation in the speaker embedding space of the ZS-TTS system. By incorporating a consistency loss, LF can be seamlessly integrated into existing ZS-TTS systems without the need for additional training stages. Experimental results show that LF significantly improves speaker similarity while preserving speech quality.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationSpeech Synthesistext-to-speechText to SpeechSimilar Papers 제목 키워드 기반
Unsupervised Panoptic Interpretation of Latent Spaces in GANs Using Space-Filling Vector Quantization
Generative adversarial networks (GANs) learn a latent space whose samples can be mapped to real-world images. Such latent spaces are difficult to interpret. Some earlier supervised methods aim to create an interpretable …
Data AugmentationQuantizationTowards Composable Distributions of Latent Space Augmentations
We propose a composable framework for latent space image augmentation that allows for easy combination of multiple augmentations. Image augmentation has been shown to be an effective technique for improving the performan…
Image Augmentationimage-classificationImage ClassificationLatent Space Bayesian Optimization with Latent Data Augmentation for Enhanced Exploration
Latent Space Bayesian Optimization (LSBO) combines generative models, typically Variational Autoencoders (VAE), with Bayesian Optimization (BO) to generate de-novo objects of interest. However, LSBO faces challenges due …
Bayesian OptimizationData AugmentationImage GenerationGenerative De-Quantization for Neural Speech Codec via Latent Diffusion
In low-bitrate speech coding, end-to-end speech coding networks aim to learn compact yet expressive features and a powerful decoder in a single network. A challenging problem as such results in unwelcome complexity incre…
DecoderQuantizationRepresentation LearningDual-Space Augmented Intrinsic-LoRA for Wind Turbine Segmentation
Accurate segmentation of wind turbine blade (WTB) images is critical for effective assessments, as it directly influences the performance of automated damage detection systems. Despite advancements in large universal vis…
Image SegmentationSegmentationSemantic Segmentation