paper-with-me

Papers

TM2D: Bimodality Driven 3D Dance Generation via Music-Text Integration

2023-04-05 · ICCV 2023 1 · Kehong Gong, Dongze Lian, Heng Chang, Chuan Guo, Zihang Jiang, Xinxin Zuo, Michael Bi Mi, Xinchao Wang

We propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce richer dance movements guided by the instructive information provided by the text. However, the lack of paired motion data with both music and text modalities limits the ability to generate dance movements that integrate both. To alleviate this challenge, we propose to utilize a 3D human motion VQ-VAE to project the motions of the two datasets into a latent space consisting of quantized vectors, which effectively mix the motion tokens from the two datasets with different distributions for training. Additionally, we propose a cross-modal transformer to integrate text instructions into motion generation architecture for generating 3D dance movements without degrading the performance of music-conditioned dance generation. To better evaluate the quality of the generated motion, we introduce two novel metrics, namely Motion Prediction Distance (MPD) and Freezing Score (FS), to measure the coherence and freezing percentage of the generated motion. Extensive experiments show that our approach can generate realistic and coherent dance movements conditioned on both text and music while maintaining comparable performance with the two single modalities. Code is available at https://garfield-kh.github.io/TM2D/.

📄 PDF Abstract BibTeX arXiv:2304.02419

Code (1)

Garfield-kh/TM2D 공식 구현 pytorch

Tasks

Motion Generationmotion predictionMotion Synthesis

Methods 이 논문이 사용한 방법론

VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…

Similar Papers 제목 키워드 기반

Every Image Listens, Every Image Dances: Music-Driven Image Animation

2025-01-30 · Zhikang Dong, Weituo Hao, Ju-Chiang Wang, Peng Zhang 외

Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video generation guided by text, music-driven d…

Image AnimationVideo Generation

GCDance: Genre-Controlled 3D Full Body Dance Generation Driven By Music

2025-02-25 · Xinran Liu, Xu Dong, Diptesh Kanojia, Wenwu Wang 외

Generating high-quality full-body dance sequences from music is a challenging task as it requires strict adherence to genre-specific choreography. Moreover, the generated sequences must be both physically realistic and p…

Rhythm

TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation

2026-04-18 · Xinran Liu, Diptesh Kanojia, Wenwu Wang, Zhenhua Feng arxiv

Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making it difficult to guide specific movements …

Cross-Modal Retrieval

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data

2026-06-29 · Kaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng 외 arxiv

Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Despite recent progress, existing methods still face two key li…

Video Generation

MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generation

2025-12-20 · Kaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng 외 arxiv

With the rise of online dance-video platforms and rapid advances in AI-generated content (AIGC), music-driven dance generation has emerged as a compelling research direction. Despite substantial progress in related domai…

Video Generation