paper-with-me

홈 › Papers

TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba

2026-03-28 · Ziyue Yang, Kaixing Yang, Xulong Tang arxiv

Music-to-dance generation has broad applications in virtual reality, dance education, and digital character animation. However, the limited coverage of existing 3D dance datasets confines current models to a narrow subset of music styles and choreographic patterns, resulting in poor generalization to real-world music. Consequently, generated dances often become overly simplistic and repetitive, substantially degrading expressiveness and realism. To tackle this problem, we present TokenDance, a two-stage music-to-dance generation framework that explicitly addresses this limitation through dual-modality tokenization and efficient token-level generation. In the first stage, we discretize both dance and music using Finite Scalar Quantization, where dance motions are factorized into upper and lower-body components with kinematic-dynamic constraints, and music is decomposed into semantic and acoustic features with dedicated codebooks to capture choreography-specific structures. In the second stage, we introduce a Local-Global-Local token-to-token generator built on a Bidirectional Mamba backbone, enabling coherent motion synthesis, strong music-dance alignment, and efficient non-autoregressive inference. Extensive experiments demonstrate that TokenDance achieves overall state-of-the-art (SOTA) performance in both generation quality and inference speed, highlighting its effectiveness and practical value for real-world music-to-dance applications.

📄 PDF Abstract BibTeX arXiv:2603.27314

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Synthesis

Similar Papers 제목 키워드 기반

DuetGen: Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling

2025-06-23 · Anindita Ghosh, Bing Zhou, Rishabh Dabral, Jian Wang 외

We present DuetGen, a novel framework for generating interactive two-person dances from music. The key challenge of this task lies in the inherent complexities of two-person dance interactions, where the partners need to…

Motion Synthesis

Global Position Aware Group Choreography using Large Language Model

2025-03-12 · Haozhou Pang, Tianwei Ding, Lanshan He, Qi Gan

Dance serves as a profound and universal expression of human culture, conveying emotions and stories through movements synchronized with music. Although some current works have achieved satisfactory results in the task o…

Language ModelingLanguage ModellingLarge Language ModelPosition

X-Dancer: Expressive Music to Human Dance Video Generation

2025-02-24 · Zeyuan Chen, Hongyi Xu, Guoxian Song, You Xie 외

We present X-Dancer, a novel zero-shot music-driven image animation pipeline that creates diverse and long-range lifelike human dance videos from a single static image. As its core, we introduce a unified transformer-dif…

Image AnimationVideo Generation

Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation

2025-12-12 · Yan Zhang, Han Zou, Lincong Feng, Cong Xie 외 arxiv

Recent pose-to-video models can translate 2D pose sequences into photorealistic, identity-preserving dance videos, so the key challenge is to generate temporally coherent, rhythm-aligned 2D poses from music, especially u…

Image Generation

SalsaAgent: A multimodal embodied language model for interactive dance generation

2026-05-28 · Payam Jome Yazdian, Zoe Stanley, Angelica Lim arxiv

Interaction between humanoids involves bidirectional and nonverbal reactivity, coordination and synchrony. Toward socially aware robots and interactive virtual agents, we present SalsaAgent, a language model that generat…