paper-with-me

홈 › Papers

EDGE: Editable Dance Generation From Music

2022-11-19 · CVPR 2023 1 · Jonathan Tseng, Rodrigo Castellon, C. Karen Liu

Dance is an important human art form, but creating new dances can be difficult and time-consuming. In this work, we introduce Editable Dance GEneration (EDGE), a state-of-the-art method for editable dance generation that is capable of creating realistic, physically-plausible dances while remaining faithful to the input music. EDGE uses a transformer-based diffusion model paired with Jukebox, a strong music feature extractor, and confers powerful editing capabilities well-suited to dance, including joint-wise conditioning, and in-betweening. We introduce a new metric for physical plausibility, and evaluate dance quality generated by our method extensively through (1) multiple quantitative metrics on physical plausibility, beat alignment, and diversity benchmarks, and more importantly, (2) a large-scale user study, demonstrating a significant improvement over previous state-of-the-art methods. Qualitative samples from our model can be found at our website.

📄 PDF Abstract BibTeX arXiv:2211.10658

Code (1)

Stanford-TML/EDGE 공식 구현 pytorch

Tasks

DiversityMotion Synthesis

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dilated Convolution 설명 없음
VQ-VAE VQ-VAE is a type of variational autoencoder that uses vector quantisation to obtain a discrete latent representation. It differs from…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
Jukebox 설명 없음

Similar Papers 제목 키워드 기반

DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions

2025-08-24 · Hengyuan Zhang, Zhe Li, Xingqun Qi, Mengze Li 외 arxiv

Generating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling use…

Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation

2026-06-22 · Seong Jong Yoo, Siyuan Peng, Felix Gu, Stratis Aloimonos 외 arxiv

Choreographic motion generation poses unique challenges for AI, demanding precise semantic control over complex, temporally structured, and expressive full-body dynamics. While existing models can synthesize motion from …

Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation

2024-07-31 · Ziya Zhou, Yuhang Wu, Zhiyue Wu, Xinyue Zhang 외

Symbolic Music, akin to language, can be encoded in discrete symbols. Recent research has extended the application of large language models (LLMs) such as GPT-4 and Llama2 to the symbolic music domain including understan…

Libretto: Giving LLM Agents a Sense of Musical Structure

2026-06-21 · Yichen Xu arxiv

Generative music systems can now produce impressive audio from text prompts, but audio outputs are difficult to inspect, edit, and diagnose as musical structure. We introduce Libretto, an agent-facing framework for symbo…

Music Generation

Music- and Lyrics-driven Dance Synthesis

2023-09-30 · Wenjie Yin, Qingyuan Yao, Yi Yu, Hang Yin 외

Lyrics often convey information about the songs that are beyond the auditory dimension, enriching the semantic meaning of movements and musical themes. Such insights are important in the dance choreography domain. Howeve…

Triplet