paper-with-me

홈 › Papers

Semi-supervised Learning for Singing Synthesis Timbre

2020-11-05 · Jordi Bonada, Merlijn Blaauw

We propose a semi-supervised singing synthesizer, which is able to learn new voices from audio data only, without any annotations such as phonetic segmentation. Our system is an encoder-decoder model with two encoders, linguistic and acoustic, and one (acoustic) decoder. In a first step, the system is trained in a supervised manner, using a labelled multi-singer dataset. Here, we ensure that the embeddings produced by both encoders are similar, so that we can later use the model with either acoustic or linguistic input features. To learn a new voice in an unsupervised manner, the pretrained acoustic encoder is used to train a decoder for the target singer. Finally, at inference, the pretrained linguistic encoder is used together with the decoder of the new voice, to produce acoustic features from linguistic input. We evaluate our system with a listening test and show that the results are comparable to those obtained with an equivalent supervised approach.

📄 PDF Abstract BibTeX arXiv:2011.02809

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

MR-SVS: Singing Voice Synthesis with Multi-Reference Encoder

2022-01-11 · Shoutong Wang, Jinglin Liu, Yi Ren, Zhen Wang 외

Multi-speaker singing voice synthesis is to generate the singing voice sung by different speakers. To generalize to new speakers, previous zero-shot singing adaptation methods obtain the timbre of the target speaker with…

Singing Voice Synthesis

Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis

2022-06-23 · Tae-Woo Kim, Min-Su Kang, Gyeong-Hoon Lee

Recently, deep learning-based generative models have been introduced to generate singing voices. One approach is to predict the parametric vocoder features consisting of explicit speech parameters. This approach has the …

Generative Adversarial NetworkMulti-Task LearningSinging Voice Synthesis

YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases

2025-12-04 · Gongyu Chen, Xiaoyu Zhang, Zhenqiang Weng, Junjie Zheng 외 arxiv

Singing voice conversion (SVC) aims to render the target singer's timbre while preserving melody and lyrics. However, existing zero-shot SVC systems remain fragile in real songs due to harmony interference, F0 errors, an…

Reinforcement LearningVoice Conversion

Self-Supervised Contrastive Learning for Singing Voices

2022-04-26 · IEEE/ACM Transactions on Audio, Speech, and Language Processing 2022 4 · Hiromu Yakura, Kento Watanabe, Masataka Goto

This study introduces self-supervised contrastive learning to acquire feature representations of singing voices. To acquire robust representations in an unsupervised manner, regular self-supervised contrastive learning t…

Contrastive LearningSinger IdentificationVocal technique classification

Learning Singing From Speech

2019-12-20 · Liqiang Zhang, Chengzhu Yu, Heng Lu, Chao Weng 외

We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples. The proposed algorithm first integrate speech and singing synthesis into a unifi…

Speech SynthesisVoice Conversion