paper-with-me

홈 › Papers

AI Choreographer: Music Conditioned 3D Dance Generation with AIST++

2021-01-21 · ICCV 2021 10 · RuiLong Li, Shan Yang, David A. Ross, Angjoo Kanazawa

We present AIST++, a new multi-modal dataset of 3D dance motion and music, along with FACT, a Full-Attention Cross-modal Transformer network for generating 3D dance motion conditioned on music. The proposed AIST++ dataset contains 5.2 hours of 3D dance motion in 1408 sequences, covering 10 dance genres with multi-view videos with known camera poses -- the largest dataset of this kind to our knowledge. We show that naively applying sequence models such as transformers to this dataset for the task of music conditioned 3D motion generation does not produce satisfactory 3D motion that is well correlated with the input music. We overcome these shortcomings by introducing key changes in its architecture design and supervision: FACT model involves a deep cross-modal transformer block with full-attention that is trained to predict $N$ future motions. We empirically show that these changes are key factors in generating long sequences of realistic dance motion that are well-attuned to the input music. We conduct extensive experiments on AIST++ with user studies, where our method outperforms recent state-of-the-art methods both qualitatively and quantitatively.

📄 PDF Abstract BibTeX arXiv:2101.08779

Code (1)

google-research/mint 공식 구현 tf

Tasks

Motion GenerationMotion SynthesisPose Estimation

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Residual Connection 설명 없음

Similar Papers 제목 키워드 기반

A Brand New Dance Partner: Music-Conditioned Pluralistic Dancing Controlled by Multiple Dance Genres

2022-01-01 · CVPR 2022 1 · Jinwoo Kim, Heeseok Oh, Seongjean Kim, Hoseok Tong 외

When coming up with phrases of movement, choreographers all have their habits as they are used to their skilled dance genres. Therefore, they tend to return certain patterns of the dance genres that they are familiar…

Generative Adversarial NetworkMotion Synthesis

DanceChat: Large Language Model-Guided Music-to-Dance Generation

2025-06-12 · Qing Wang, Xiaohang Yang, Yilan Dong, Naveen Raj Govindaraj 외

Music-to-dance generation aims to synthesize human dance motion conditioned on musical input. Despite recent progress, significant challenges remain due to the semantic gap between music and dance motion, as music offers…

Language ModelingLanguage ModellingLarge Language ModelMotion Synthesis+1

LM2D: Lyrics- and Music-Driven Dance Synthesis

2024-03-14 · Wenjie Yin, Xuejiao Zhao, Yi Yu, Hang Yin 외

Dance typically involves professional choreography with complex movements that follow a musical rhythm and can also be influenced by lyrical content. The integration of lyrics in addition to the auditory dimension, enric…

Motion GenerationPose EstimationRhythm

MIDGET: Music Conditioned 3D Dance Generation

2024-04-18 · Jinwu Wang, Wei Mao, Miaomiao Liu

In this paper, we introduce a MusIc conditioned 3D Dance GEneraTion model, named MIDGET based on Dance motion Vector Quantised Variational AutoEncoder (VQ-VAE) model and Motion Generative Pre-Training (GPT) model to gene…

Rhythm

GCDance: Genre-Controlled 3D Full Body Dance Generation Driven By Music

2025-02-25 · Xinran Liu, Xu Dong, Diptesh Kanojia, Wenwu Wang 외

Generating high-quality full-body dance sequences from music is a challenging task as it requires strict adherence to genre-specific choreography. Moreover, the generated sequences must be both physically realistic and p…

Rhythm