PhaseAug: A Differentiable Augmentation for Speech Synthesis to Simulate One-to-Many Mapping
Previous generative adversarial network (GAN)-based neural vocoders are trained to reconstruct the exact ground truth waveform from the paired mel-spectrogram and do not consider the one-to-many relationship of speech synthesis. This conventional training causes overfitting for both the discriminators and the generator, leading to the periodicity artifacts in the generated audio signal. In this work, we present PhaseAug, the first differentiable augmentation for speech synthesis that rotates the phase of each frequency bin to simulate one-to-many mapping. With our proposed method, we outperform baselines without any architecture modification. Code and audio samples will be available at https://github.com/mindslab-ai/phaseaug.
Code (2)
Tasks
Generative Adversarial NetworkSpeech SynthesisSimilar Papers 제목 키워드 기반
VIFS: An End-to-End Variational Inference for Foley Sound Synthesis
The goal of DCASE 2023 Challenge Task 7 is to generate various sound clips for Foley sound synthesis (FSS) by "category-to-sound" approach. "Category" is expressed by a single index while corresponding "sound" covers div…
Speech Synthesistext-to-speechText to SpeechVariational InferenceAudio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
Automatic detection of synthetic speech is becoming increasingly important as current synthesis methods are both near indistinguishable from human speech and widely accessible to the public. Audio watermarking and other …
Face SwappingSpeech SynthesisAutoTTS: End-to-End Text-to-Speech Synthesis through Differentiable Duration Modeling
Parallel text-to-speech (TTS) models have recently enabled fast and highly-natural speech synthesis. However, they typically require external alignment models, which are not necessarily optimized for the decoder as they …
DecoderSpeech Synthesistext-to-speechText to Speech+1Embedding a Differentiable Mel-cepstral Synthesis Filter to a Neural Speech Synthesis System
This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral synthesis filter is explicitly embedded …
GPUSpeech SynthesisSpeech Synthesis as Augmentation for Low-Resource ASR
Speech synthesis might hold the key to low-resource speech recognition. Data augmentation techniques have become an essential part of modern speech recognition training. Yet, they are simple, naive, and rarely reflect re…
Data Augmentationspeech-recognitionSpeech RecognitionSpeech Synthesis