paper-with-me

Papers

SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning

2026-08-17 · Tao Feng, Xu Li, Xiangyang Luo, Ming Wen, Huadai Liu, Chen Zhang, Wei Xue arxiv

Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate the vocals. Existing music-conditioned methods focus primarily on choreography, while speech-driven models generally assume that the visible subject produces the input voice, leaving this combined setting largely underexplored. We introduce SingDance, a unified video diffusion framework that formulates controllable vocal articulation as a semantic role: the visible subject is either the source, who produces the vocal signal, or the listener, who receives it from an off-screen performer. Hard-compact routing selects task-relevant speech, music, and role conditions, which are composed through frame-wise joint audio injection; source and listener retain the same speech pathway. Training uses asymmetric supervision: on-screen speaking and curated off-screen conversational-response videos establish role control, while instrumental and song-based dancing-only videos establish music-conditioned body motion. The target Song/Source configuration is never observed during training. At inference, assigning the source role to a song composes separately learned articulation and song-conditioned dance capabilities, enabling compositional zero-shot singing-and-dancing. Experiments demonstrate strong motion--beat alignment and visual fidelity, reliable paired switching of vocal articulation while preserving music-aligned body motion, and highly competitive lip synchronization with substantially fewer generation-time parameters than the strongest speech-driven baseline evaluated.

📄 PDF Abstract BibTeX arXiv:2608.16220

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control

2024-09-24 · Yu Zhang, Ziyue Jiang, RuiQi Li, Changhao Pan 외

Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunc…

ClusteringLanguage ModellingQuantizationRhythm+2

SaMoye: Zero-shot Singing Voice Conversion Model Based on Feature Disentanglement and Enhancement

2024-07-10 · ZiHao Wang, Le Ma, Yongsheng Feng, Xin Pan 외

Singing voice conversion (SVC) aims to convert a singer's voice to another singer's from a reference audio while keeping the original semantics. However, existing SVC methods can hardly perform zero-shot due to incomplet…

DisentanglementVoice Conversion

RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking

2025-09-25 · Zhenguo Sun, Yibo Peng, Yuan Meng, Xukun Li 외 arxiv

Long-horizon, high-dynamic motion tracking on humanoids remains brittle because absolute joint commands cannot compensate model-plant mismatch, leading to error accumulation. We propose RobotDancing, a simple, scalable f…

Reinforcement Learning

YingMusic-SVC: Real-World Robust Zero-Shot Singing Voice Conversion with Flow-GRPO and Singing-Specific Inductive Biases

2025-12-04 · Gongyu Chen, Xiaoyu Zhang, Zhenqiang Weng, Junjie Zheng 외 arxiv

Singing voice conversion (SVC) aims to render the target singer's timbre while preserving melody and lyrics. However, existing zero-shot SVC systems remain fragile in real songs due to harmony interference, F0 errors, an…

Reinforcement LearningVoice Conversion

TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis

2025-05-20 · Yu Zhang, Wenxiang Guo, Changhao Pan, Dongyu Yao 외

Customizable multilingual zero-shot singing voice synthesis (SVS) has various potential applications in music composition and short video dubbing. However, existing SVS models overly depend on phoneme and note boundary a…

Contrastive LearningSinging Voice SynthesisStyle Transfer