A Survey on Recent Deep Learning-driven Singing Voice Synthesis Systems
Singing voice synthesis (SVS) is a task that aims to generate audio signals according to musical scores and lyrics. With its multifaceted nature concerning music and language, producing singing voices indistinguishable from that of human singers has always remained an unfulfilled pursuit. Nonetheless, the advancements of deep learning techniques have brought about a substantial leap in the quality and naturalness of synthesized singing voice. This paper aims to review some of the state-of-the-art deep learning-driven SVS systems. We intend to summarize their deployed model architectures and identify the strengths and limitations for each of the introduced systems. Thereby, we picture the recent advancement trajectory of this field and conclude the challenges left to be resolved both in commercial applications and academic research.
Code (0)
등록된 구현이 없습니다.
Tasks
Deep LearningSinging Voice SynthesisSimilar Papers 제목 키워드 기반
Controllable Singing Voice Synthesis using Phoneme-Level Energy Sequence
Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise con…
ConSinger: Efficient High-Fidelity Singing Voice Generation with Minimal Steps
Singing voice synthesis (SVS) system is expected to generate high-fidelity singing voice from given music scores (lyrics, duration and pitch). Recently, diffusion models have performed well in this field. However, sacrif…
Singing Voice SynthesisA Melody-Unsupervision Model for Singing Voice Synthesis
Recent studies in singing voice synthesis have achieved high-quality results leveraging advances in text-to-speech models based on deep neural networks. One of the main issues in training singing voice synthesis models i…
modelSinging Voice Synthesistext-to-speechText to SpeechMLP Singer: Towards Rapid Parallel Singing Voice Synthesis
Recent developments in deep learning have significantly improved the quality of synthesized singing voice audio. However, prominent neural singing voice synthesis systems suffer from slow inference speed due to their aut…
image-classificationSinging Voice SynthesisTowards Improving the Expressiveness of Singing Voice Synthesis with BERT Derived Semantic Information
This paper presents an end-to-end high-quality singing voice synthesis (SVS) system that uses bidirectional encoder representation from Transformers (BERT) derived semantic embeddings to improve the expressiveness of the…
Singing Voice Synthesis