paper-with-me

홈 › Papers

Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment

2024-04-14 · Zhiqing Hong, Rongjie Huang, Xize Cheng, Yongqi Wang, RuiQi Li, Fuming You, Zhou Zhao, Zhimeng Zhang

A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, we propose a novel task called text-to-song synthesis which incorporating both vocals and accompaniments generation. We develop Melodist, a two-stage text-to-song method that consists of singing voice synthesis (SVS) and vocal-to-accompaniment (V2A) synthesis. Melodist leverages tri-tower contrastive pretraining to learn more effective text representation for controllable V2A synthesis. A Chinese song dataset mined from a music website is built up to alleviate data scarcity for our research. The evaluation results on our dataset demonstrate that Melodist can synthesize songs with comparable quality and style consistency. Audio samples can be found in https://text2songMelodist.github.io/Sample/.

📄 PDF Abstract BibTeX arXiv:2404.09313

Code (0)

등록된 구현이 없습니다.

Tasks

Music GenerationSinging Voice Synthesis

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models

2024-05-16 · Ziyu Wang, Lejun Min, Gus Xia

Recent deep music generation studies have put much emphasis on long-term generation with structures. However, we are yet to see high-quality, well-structured whole-song generation. In this paper, we make the first attemp…

Music Generation

DiffRhythm+: Controllable and Flexible Full-Length Song Generation with Preference Optimization

2025-07-17 · Huakang Chen, Yuepeng Jiang, Guobin Ma, Chunbo Hao 외

Songs, as a central form of musical art, exemplify the richness of human intelligence and creativity. While recent advances in generative modeling have enabled notable progress in long-form song generation, current syste…

Descriptive

Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control

2026-01-07 · Changhao Jiang, Jiahao Chen, Zhenghao Xiang, Zhixiong Yang 외 arxiv

Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, while academic research remains largely non-reproducible due to the lack of publicly available training data, hindering…

CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls

2024-12-13 · Li Chai, Donglin Wang

Lyric-to-melody generation is a highly challenging task in the field of AI music generation. Due to the difficulty of learning strict yet weak correlations between lyrics and melodies, previous methods have suffered from…

DecoderMusic GenerationSentence

Sing-On-Your-Beat: Simple Text-Controllable Accompaniment Generations

2024-11-03 · Quoc-Huy Trinh, Minh-Van Nguyen, Trong-Hieu Nguyen Mau, Khoa Tran 외

Singing is one of the most cherished forms of human entertainment. However, creating a beautiful song requires an accompaniment that complements the vocals and aligns well with the song instruments and genre. With advanc…