paper-with-me

홈 › Papers

MR-SVS: Singing Voice Synthesis with Multi-Reference Encoder

2022-01-11 · Shoutong Wang, Jinglin Liu, Yi Ren, Zhen Wang, Changliang Xu, Zhou Zhao

Multi-speaker singing voice synthesis is to generate the singing voice sung by different speakers. To generalize to new speakers, previous zero-shot singing adaptation methods obtain the timbre of the target speaker with a fixed-size embedding from single reference audio. However, they face several challenges: 1) the fixed-size speaker embedding is not powerful enough to capture full details of the target timbre; 2) single reference audio does not contain sufficient timbre information of the target speaker; 3) the pitch inconsistency between different speakers also leads to a degradation in the generated voice. In this paper, we propose a new model called MR-SVS to tackle these problems. Specifically, we employ both a multi-reference encoder and a fixed-size encoder to encode the timbre of the target speaker from multiple reference audios. The Multi-reference encoder can capture more details and variations of the target timbre. Besides, we propose a well-designed pitch shift method to address the pitch inconsistency problem. Experiments indicate that our method outperforms the baseline method both in naturalness and similarity.

📄 PDF Abstract BibTeX arXiv:2201.03864

Code (0)

등록된 구현이 없습니다.

Tasks

Singing Voice Synthesis

Similar Papers 제목 키워드 기반

An Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures

2021-08-06 · Dengfeng Ke, Yuxing Lu, Xudong Liu, Yanyan Xu 외

With the rapid development of neural network architectures and speech processing models, singing voice synthesis with neural networks is becoming the cutting-edge technique of digital music production. In this work, in o…

DecoderSinging Voice Synthesis

DeepSinger: Singing Voice Synthesis with Data Mined From the Web

2020-07-09 · Yi Ren, Xu Tan, Tao Qin, Jian Luan 외

In this paper, we develop DeepSinger, a multi-lingual multi-singer singing voice synthesis (SVS) system, which is built from scratch using singing training data mined from music websites. The pipeline of DeepSinger consi…

SentenceSinging Voice Synthesis

StyleSinger: Style Transfer for Out-of-Domain Singing Voice Synthesis

2023-12-17 · Yu Zhang, Rongjie Huang, RuiQi Li, Jinzheng He 외

Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from ref…

QuantizationSinging Voice SynthesisStyle Transfer

TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control

2024-09-24 · Yu Zhang, Ziyue Jiang, RuiQi Li, Changhao Pan 외

Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunc…

ClusteringLanguage ModellingQuantizationRhythm+2

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

2025-05-20 · Yu Zhang, Wenxiang Guo, Changhao Pan, Dongyu Yao 외

Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary a…

Contrastive LearningSinging Voice SynthesisStyle Transfer