paper-with-me

홈 › Papers

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

2026-07-10 · Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Bang Zhang arxiv

Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.

📄 PDF Abstract BibTeX arXiv:2607.09581

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation

2025-11-24 · Jiaming Zhang, Shengming Cao, Rui Li, Xiaotong Zhao 외 arxiv

Preserving first-frame identity while ensuring precise motion control is a fundamental challenge in human image animation. The Image-to-Motion Binding process of the dominant Reference-to-Video (R2V) paradigm overlooks c…

TCDiff++: An End-to-end Trajectory-Controllable Diffusion Model for Harmonious Music-Driven Group Choreography

2025-06-23 · Yuqin Dai, Wanlu Zhu, Ronghui Li, Xiu Li 외

Music-driven dance generation has garnered significant attention due to its wide range of industrial applications, particularly in the creation of group choreography. During the group dance generation process, however, m…

Harmonious Group Choreography with Trajectory-Controllable Diffusion

2024-03-10 · Yuqin Dai, Wanlu Zhu, Ronghui Li, Zeping Ren 외

Creating group choreography from music has gained attention in cultural entertainment and virtual reality, aiming to coordinate visually cohesive and diverse group movements. Despite increasing interest, recent works fac…

Motion Synthesis

ST-GDance: Long-Term and Collision-Free Group Choreography from Music

2025-07-29 · Jing Xu, Weiqiang Wang, Cunjian Chen, Jun Liu 외 arxiv

Group dance generation from music has broad applications in film, gaming, and animation production. However, it requires synchronizing multiple dancers while maintaining spatial coordination. As the number of dancers and…

Music-Driven Group Choreography

2023-03-22 · CVPR 2023 1 · Nhat Le, Thang Pham, Tuong Do, Erman Tjiputra 외

Music-driven choreography is a challenging problem with a wide variety of industrial applications. Recently, many methods have been proposed to synthesize dance motions from music for a single dancer. However, generating…

Motion Synthesis