paper-with-me

홈 › Papers

SegTune: Structured and Fine-Grained Control for Song Generation

2026-05-31 · Yuejiao Wang, Zihao Ji, Pengfei Cai, Xu Li, Haorui Zheng, Zewen Song, Zhongliang Liu, Chen Zhang, Pengfei Wan arxiv

Recent advances in neural song generation have enabled high-quality synthesis from lyrics and global textual prompts. However, most systems fail to model temporally varying attributes of songs, severely limiting fine-grained control over musical structure and dynamics. To address this, we propose SegTune, a Diffusion Transformer-based framework enabling structured and fine-grained controllability by allowing users or large language models (LLMs) to specify local musical descriptions aligned to song segments. These segment prompts are temporally broadcast to corresponding time windows, while global prompts ensure stylistic coherence. To support precise lyric-to-music alignment, we introduce an LLM-based duration predictor that autoregressively generates sentence-level timestamps in LyRiCs format. We further construct a large-scale data pipeline for high-quality song collection with aligned lyrics and prompts, and propose new metrics to evaluate segment alignment and vocal consistency. Experiments demonstrate that SegTune outperforms existing baselines in both musicality and controllability. Visit our project page (https://github.com/KlingAIResearch/SegTune) for codes and more generated songs.

📄 PDF Abstract BibTeX arXiv:2606.02638

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls

2024-12-13 · Li Chai, Donglin Wang

Lyric-to-melody generation is a highly challenging task in the field of AI music generation. Due to the difficulty of learning strict yet weak correlations between lyrics and melodies, previous methods have suffered from…

DecoderMusic GenerationSentence

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

2025-07-28 · Renhang Liu, Chia-Yu Hung, Navonil Majumder, Taylor Gautreaux 외 arxiv

Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and faithful audio outputs capturing to speech a…

Audio Generation

Muse: Towards Reproducible Long-Form Song Generation with Fine-Grained Style Control

2026-01-07 · Changhao Jiang, Jiahao Chen, Zhenghao Xiang, Zhixiong Yang 외 arxiv

Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, while academic research remains largely non-reproducible due to the lack of publicly available training data, hindering…

SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment

2026-04-16 · Dapeng Wu, Shun Lei, Wei Tan, Guangzheng Li 외 arxiv

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In th…

SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation

2025-02-18 · Zihan Liu, Shuangrui Ding, Zhixiong Zhang, Xiaoyi Dong 외

Text-to-song generation, the task of creating vocals and accompaniment from textual inputs, poses significant challenges due to domain complexity and data scarcity. Existing approaches often employ multi-stage generation…

Voice Cloning