paper-with-me

홈 › Papers

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance

2025-09-24 · Junchuan Zhao, Wei Zeng, Tianle Lyu, Ye Wang arxiv

Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete codec-based speech synthesis has enabled zero-shot generation via in-context learning, directly extending these techniques to SVS remains non-trivial due to the requirement for precise melody control. In particular, prompt-based generation often introduces prosody leakage, where pitch information is inadvertently entangled within the timbre prompt, compromising controllability. We present CoMelSinger, a zero-shot SVS framework that enables structured and disentangled melody control within a discrete codec modeling paradigm. Built on the non-autoregressive MaskGCT architecture, CoMelSinger replaces conventional text inputs with lyric and pitch tokens, preserving in-context generalization while enhancing melody conditioning. To suppress prosody leakage, we propose a coarse-to-fine contrastive learning strategy that explicitly regularizes pitch redundancy between the acoustic prompt and melody input. Furthermore, we incorporate a lightweight encoder-only Singing Voice Transcription (SVT) module to align acoustic tokens with pitch and duration, offering fine-grained frame-level supervision. Experimental results demonstrate that CoMelSinger achieves notable improvements in pitch accuracy, timbre consistency, and zero-shot transferability over competitive baselines. Audio samples are available at https://danny-nus.github.io/CoMelSinger/.

📄 PDF Abstract BibTeX arXiv:2509.19883

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningSpeech Synthesis

Similar Papers 제목 키워드 기반

NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

2023-04-18 · Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu 외

Scaling text-to-speech (TTS) to large-scale, multi-speaker, and in-the-wild datasets is important to capture the diversity in human speech such as speaker identities, prosodies, and styles (e.g., singing). Current large …

In-Context LearningSpeech Synthesistext-to-speechText to Speech

Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion

2024-06-04 · RuiQi Li, Rongjie Huang, Yongqi Wang, Zhiqing Hong 외

Speech-to-singing voice conversion (STS) task always suffers from data scarcity, because it requires paired speech and singing data. Compounding this issue are the challenges of content-pitch alignment and the suboptimal…

In-Context LearningLanguage ModelingLanguage ModellingRhythm+3

TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control

2024-09-24 · Yu Zhang, Ziyue Jiang, RuiQi Li, Changhao Pan 외

Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunc…

ClusteringLanguage ModellingQuantizationRhythm+2

SaMoye: Zero-shot Singing Voice Conversion Model Based on Feature Disentanglement and Enhancement

2024-07-10 · ZiHao Wang, Le Ma, Yongsheng Feng, Xin Pan 외

Singing voice conversion (SVC) aims to convert a singer's voice to another singer's from a reference audio while keeping the original semantics. However, existing SVC methods can hardly perform zero-shot due to incomplet…

DisentanglementVoice Conversion

YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases

2025-12-04 · Gongyu Chen, Xiaoyu Zhang, Zhenqiang Weng, Junjie Zheng 외 arxiv

Singing voice conversion (SVC) aims to render the target singer's timbre while preserving melody and lyrics. However, existing zero-shot SVC systems remain fragile in real songs due to harmony interference, F0 errors, an…

Reinforcement LearningVoice Conversion