paper-with-me

Papers

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

2026-06-29 · Shun Lei, Huaicheng Zhang, Dapeng Wu, Yaoxun Xu, Lishi Zuo, Wei Tan, Hangting Chen, Guangzheng Li, Jianwei Yu, Zhiyong Wu, Dong Yu arxiv

Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics and prompts. Existing language model-based systems face a structural trade-off: mixed-token modeling preserves vocal-instrument coordination but obscures track-specific details, whereas dual-track prediction improves acoustics but requires longer sequences and weakens global planning. We present LeVo 2, a hybrid LLM-Diffusion framework for controllable full-length song generation. LeVo 2 formulates this trade-off as hierarchical modeling: LeLM first predicts mixed tokens for semantic planning, then predicts vocal and accompaniment tokens in parallel for track-specific refinement, while a diffusion-based Music Codec reconstructs full-length waveforms. A central contribution of this extended version is an aesthetics-guided training schedule for alignment. During pre-training, an automated music aesthetic evaluation framework assigns musicality-tier conditions to large-scale data, providing musicality priors before preference alignment. Progressive post-training applies SFT, large-scale offline DPO, and closed-loop semi-online DPO to separately improve generation quality, controllability, and musicality. Modular extension then trains the Track-Specific LM for acoustic refinement while preserving the aligned semantic planner. This schedule separates musicality learning, controllability alignment, and acoustic refinement, mitigating optimization conflict and the limitations of static offline preference pairs. Expert listening tests and objective evaluations show that LeVo 2 outperforms open-source baselines across six subjective dimensions, and approaches leading commercial systems on several listening metrics. Ablations validate the effects of the training strategy, aesthetics guidance, scaling, and hierarchical architecture.

📄 PDF Abstract BibTeX arXiv:2606.30642

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LeVo: High-Quality Song Generation with Multi-Preference Alignment

2025-06-09 · Shun Lei, Yaoxun Xu, Zhiwei Lin, Huaicheng Zhang 외

Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing approaches still struggle with the comp…

Instruction FollowingMusic Generation

Neural Melody Composition from Lyrics

2018-09-12 · Hangbo Bao, Shaohan Huang, Furu Wei, Lei Cui 외

In this paper, we study a novel task that learns to compose music from natural language. Given the lyrics as input, we propose a melody composition model that generates lyrics-conditional melody as well as the exact alig…

Decoder

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment

2025-07-28 · Renhang Liu, Chia-Yu Hung, Navonil Majumder, Taylor Gautreaux 외 arxiv

Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and faithful audio outputs capturing to speech a…

Audio Generation

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

2026-06-04 · Shangheng Du, Xiangchao Yan, Jinxin Shi, Zongsheng Cao 외 arxiv

Large language model (LLM) agents are increasingly applied to long-horizon tasks such as scientific discovery and machine learning engineering (MLE), where sustained self-evolution becomes a key capability. However, exis…

Domain GeneralizationCode Generation

Multifunctional pH sensitive 3D scaffolds for treatment and prevention of bone infection

2021-04-19 · Cicuendez M, Doadrio JC, Hernandez A, Portoles MT 외

Multifunctional-therapeutic 3D scaffolds have been prepared. These biomaterials are able to destroy the S. aureus bacteria biofilm and to allow bone regeneration at the same time. The present study is focused on the desi…