paper-with-me

Papers

Versatile Framework for Song Generation with Prompt-based Control

2025-04-27 · Yu Zhang, Wenxiang Guo, Changhao Pan, Zhiyuan Zhu, RuiQi Li, Jingyu Lu, Rongjie Huang, Ruiyuan Zhang, Zhiqing Hong, Ziyue Jiang, Zhou Zhao

Song generation focuses on producing controllable high-quality songs based on various prompts. However, existing methods struggle to generate vocals and accompaniments with prompt-based control and proper alignment. Additionally, they fall short in supporting various tasks. To address these challenges, we introduce VersBand, a multi-task song generation framework for synthesizing high-quality, aligned songs with prompt-based control. VersBand comprises these primary models: 1) VocalBand, a decoupled model, leverages the flow-matching method for generating singing styles, pitches, and mel-spectrograms, allowing fast, high-quality vocal generation with style control. 2) AccompBand, a flow-based transformer model, incorporates the Band-MOE, selecting suitable experts for enhanced quality, alignment, and control. This model allows for generating controllable, high-quality accompaniments aligned with vocals. 3) Two generation models, LyricBand for lyrics and MelodyBand for melodies, contribute to the comprehensive multi-task song generation system, allowing for extensive control based on multiple prompts. Experimental results demonstrate that VersBand performs better over baseline models across multiple song generation tasks using objective and subjective metrics. Audio samples are available at https://aaronz345.github.io/VersBandDemo.

📄 PDF Abstract BibTeX arXiv:2504.19062

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SegTune: Structured and Fine-Grained Control for Song Generation

2026-05-31 · Yuejiao Wang, Zihao Ji, Pengfei Cai, Xu Li 외 arxiv

Recent advances in neural song generation have enabled high-quality synthesis from lyrics and global textual prompts. However, most systems fail to model temporally varying attributes of songs, severely limiting fine-gra…

Accompanied Singing Voice Synthesis with Fully Text-controlled Melody

2024-07-02 · RuiQi Li, Zhiqing Hong, Yongqi Wang, Lichao Zhang 외

Text-to-song (TTSong) is a music generation task that synthesizes accompanied singing voices. Current TTSong methods, inherited from singing voice synthesis (SVS), require melody-related information that can sometimes be…

Language ModelingLanguage ModellingMusic GenerationSinging Voice Synthesis

SongCreator: Lyrics-based Universal Song Generation

2024-09-09 · Shun Lei, Yixuan Zhou, Boshi Tang, Max W. Y. Lam 외

Music is an integral part of human culture, embodying human intelligence and creativity, of which songs compose an essential part. While various aspects of song generation have been explored by previous works, such as si…

Language ModellingMusic Generation

SongRewriter: A Chinese Song Rewriting System with Controllable Content and Rhyme Scheme

2022-11-28 · Yusen Sun, Liangyou Li, Qun Liu, Dit-yan Yeung

Although lyrics generation has achieved significant progress in recent years, it has limited practical applications because the generated lyrics cannot be performed without composing compatible melodies. In this work, we…

Rhythm

Sing-On-Your-Beat: Simple Text-Controllable Accompaniment Generations

2024-11-03 · Quoc-Huy Trinh, Minh-Van Nguyen, Trong-Hieu Nguyen Mau, Khoa Tran 외

Singing is one of the most cherished forms of human entertainment. However, creating a beautiful song requires an accompaniment that complements the vocals and aligns well with the song instruments and genre. With advanc…