paper-with-me

Papers

DiffDance: Cascaded Human Motion Diffusion Model for Dance Generation

2023-08-05 · Qiaosong Qi, Le Zhuo, Aixi Zhang, Yue Liao, Fei Fang, Si Liu, Shuicheng Yan

When hearing music, it is natural for people to dance to its rhythm. Automatic dance generation, however, is a challenging task due to the physical constraints of human motion and rhythmic alignment with target music. Conventional autoregressive methods introduce compounding errors during sampling and struggle to capture the long-term structure of dance sequences. To address these limitations, we present a novel cascaded motion diffusion model, DiffDance, designed for high-resolution, long-form dance generation. This model comprises a music-to-dance diffusion model and a sequence super-resolution diffusion model. To bridge the gap between music and motion for conditional generation, DiffDance employs a pretrained audio representation learning model to extract music embeddings and further align its embedding space to motion via contrastive loss. During training our cascaded diffusion model, we also incorporate multiple geometric losses to constrain the model outputs to be physically plausible and add a dynamic loss weight that adaptively changes over diffusion timesteps to facilitate sample diversity. Through comprehensive experiments performed on the benchmark dataset AIST++, we demonstrate that DiffDance is capable of generating realistic dance sequences that align effectively with the input music. These results are comparable to those achieved by state-of-the-art autoregressive methods.

📄 PDF Abstract BibTeX arXiv:2308.02915

Code (0)

등록된 구현이 없습니다.

Tasks

Representation LearningRhythmSuper-Resolution

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation

2025-12-20 · Kaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng 외 arxiv

With the rise of online dance-video platforms and rapid advances in AI-generated content (AIGC), music-driven dance generation has emerged as a compelling research direction. Despite substantial progress in related domai…

Video Generation

Generating Human Motion Videos using a Cascaded Text-to-Video Framework

2025-10-04 · Hyelin Nam, Hyojun Go, Byeongjun Park, Byung-Hoon Kim 외 arxiv

Human video generation is becoming an increasingly important task with broad applications in graphics, entertainment, and embodied AI. Despite the rapid progress of video diffusion models (VDMs), their use for general-pu…

Video Generation

Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control

2025-12-29 · Zhe Li, Cheng Chi, Yangyang Wei, Boan Zhu 외 arxiv

Humans intuitively move to sound, but current humanoid robots lack expressive improvisational capabilities, confined to predefined motions or sparse commands. Generating motion from audio and then retargeting it to robot…

Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation

2026-05-13 · Deli Cai, Haoyang Ma, Changxing Ding arxiv

Trajectory-controlled human motion generation aims to synthesize realistic human motions conditioned on both textual descriptions and spatial trajectories. However, existing methods suffer from two critical limitations: …

Bidirectional Autoregessive Diffusion Model for Dance Generation

2024-01-01 · CVPR 2024 1 · Canyu Zhang, YouBao Tang, Ning Zhang, Ruei-Sung Lin 외

Dance serves as a powerful medium for expressing human emotions but the lifelike generation of dance is still a considerable challenge. Recently diffusion models have showcased remarkable generative abilities across …

modelMotion Generation