paper-with-me

Papers

Accompanied Singing Voice Synthesis with Fully Text-controlled Melody

2024-07-02 · RuiQi Li, Zhiqing Hong, Yongqi Wang, Lichao Zhang, Rongjie Huang, Siqi Zheng, Zhou Zhao

Text-to-song (TTSong) is a music generation task that synthesizes accompanied singing voices. Current TTSong methods, inherited from singing voice synthesis (SVS), require melody-related information that can sometimes be impractical, such as music scores or MIDI sequences. We present MelodyLM, the first TTSong model that generates high-quality song pieces with fully text-controlled melodies, achieving minimal user requirements and maximum control flexibility. MelodyLM explicitly models MIDI as the intermediate melody-related feature and sequentially generates vocal tracks in a language model manner, conditioned on textual and vocal prompts. The accompaniment music is subsequently synthesized by a latent diffusion model with hybrid conditioning for temporal alignment. With minimal requirements, users only need to input lyrics and a reference voice to synthesize a song sample. For full control, just input textual prompts or even directly input MIDI. Experimental results indicate that MelodyLM achieves superior performance in terms of both objective and subjective metrics. Audio samples are available at https://melodylm666.github.io.

📄 PDF Abstract BibTeX arXiv:2407.02049

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMusic GenerationSinging Voice Synthesis

Methods 이 논문이 사용한 방법론

Latent Diffusion Model Diffusion models applied to latent spaces, which are normally built with (Variational) Autoencoders.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational Autoencoder

2024-12-12 · Jianwei Cui, Yu Gu, Shihao Chen, Jie Zhang 외

Singing Voice Synthesis (SVS) aims to generate singing voices of high fidelity and expressiveness. Conventional SVS systems usually utilize an acoustic model to transform a music score into acoustic features, followed by…

Audio SynthesisSinging Voice Synthesistext-to-speechText to Speech

A Melody-Unsupervision Model for Singing Voice Synthesis

2021-10-13 · Soonbeom Choi, Juhan Nam

Recent studies in singing voice synthesis have achieved high-quality results leveraging advances in text-to-speech models based on deep neural networks. One of the main issues in training singing voice synthesis models i…

modelSinging Voice Synthesistext-to-speechText to Speech

GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks

2024-09-20 · Yu Zhang, Changhao Pan, Wenxiang Guo, RuiQi Li 외

The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited div…

AllSinging Voice SynthesisStyle TransferVocal technique classification

An Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures

2021-08-06 · Dengfeng Ke, Yuxing Lu, Xudong Liu, Yanyan Xu 외

With the rapid development of neural network architectures and speech processing models, singing voice synthesis with neural networks is becoming the cutting-edge technique of digital music production. In this work, in o…

DecoderSinging Voice Synthesis

Learning Singing From Speech

2019-12-20 · Liqiang Zhang, Chengzhu Yu, Heng Lu, Chao Weng 외

We propose an algorithm that is capable of synthesizing high quality target speaker's singing voice given only their normal speech samples. The proposed algorithm first integrate speech and singing synthesis into a unifi…

Speech SynthesisVoice Conversion