SingAug: Data Augmentation for Singing Voice Synthesis with Cycle-consistent Training Strategy
Deep learning based singing voice synthesis (SVS) systems have been demonstrated to flexibly generate singing with better qualities, compared to conventional statistical parametric based methods. However, neural systems are generally data-hungry and have difficulty to reach reasonable singing quality with limited public available training data. In this work, we explore different data augmentation methods to boost the training of SVS systems, including several strategies customized to SVS based on pitch augmentation and mix-up augmentation. To further stabilize the training, we introduce the cycle-consistent training strategy. Extensive experiments on two public singing databases demonstrate that our proposed augmentation methods and the stabilizing training strategy can significantly improve the performance on both objective and subjective evaluations.
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationSinging Voice SynthesisSimilar Papers 제목 키워드 기반
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
Diffusion-based singing voice conversion (SVC) models have shown better synthesis quality compared to traditional methods. However, in cross-domain SVC scenarios, where there is a significant disparity in pitch between t…
SSIMVoice ConversionRapping-Singing Voice Synthesis based on Phoneme-level Prosody Control
In this paper, a text-to-rapping/singing system is introduced, which can be adapted to any speaker's voice. It utilizes a Tacotron-based multispeaker acoustic model trained on read-only speech data and which provides pro…
Singing Voice SynthesisvalidWeSinger: Data-augmented Singing Voice Synthesis with Auxiliary Losses
In this paper, we develop a new multi-singer Chinese neural singing voice synthesis (SVS) system named WeSinger. To improve the accuracy and naturalness of synthesized singing voice, we design several specifical modules …
Data AugmentationDecoderRhythmSinging Voice SynthesisAn Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures
With the rapid development of neural network architectures and speech processing models, singing voice synthesis with neural networks is becoming the cutting-edge technique of digital music production. In this work, in o…
DecoderSinging Voice SynthesisLearning Singing From Speech
We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples. The proposed algorithm first integrate speech and singing synthesis into a unifi…
Speech SynthesisVoice Conversion