paper-with-me

홈 › Papers

Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model

2025-03-28 · CVPR 2025 1 · Changchang Sun, Gaowen Liu, Charles Fleming, Yan Yan

Conditional diffusion models have gained increasing attention since their impressive results for cross-modal synthesis, where the strong alignment between conditioning input and generated output can be achieved by training a time-conditioned U-Net augmented with cross-attention mechanism. In this paper, we focus on the problem of generating music synchronized with rhythmic visual cues of the given dance video. Considering that bi-directional guidance is more beneficial for training a diffusion model, we propose to enhance the quality of generated music and its synchronization with dance videos by adopting both positive rhythmic information and negative ones (PN-Diffusion) as conditions, where a dual diffusion and reverse processes is devised. Specifically, to train a sequential multi-modal U-Net structure, PN-Diffusion consists of a noise prediction objective for positive conditioning and an additional noise prediction objective for negative conditioning. To accurately define and select both positive and negative conditioning, we ingeniously utilize temporal correlations in dance videos, capturing positive and negative rhythmic cues by playing them forward and backward, respectively. Through subjective and objective evaluations of input-output correspondence in terms of dance-music beat alignment and the quality of generated music, experimental results on the AIST++ and TikTok dance video datasets demonstrate that our model outperforms SOTA dance-to-music generation models.

📄 PDF Abstract BibTeX arXiv:2503.22138

Code (0)

등록된 구현이 없습니다.

Tasks

Music Generation

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
U-Net 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Performance Conditioning for Diffusion-Based Multi-Instrument Music Synthesis

2023-09-21 · Ben Maman, Johannes Zeitler, Meinard Müller, Amit H. Bermano

Generating multi-instrument music from symbolic music representations is an important task in Music Information Retrieval (MIR). A central but still largely unsolved problem in this context is musically and acoustically …

FADInformation RetrievalMusic Information RetrievalRetrieval

Reframing Music-Driven 2D Dance Pose Generation as Multi-Channel Image Generation

2025-12-12 · Yan Zhang, Han Zou, Lincong Feng, Cong Xie 외 arxiv

Recent pose-to-video models can translate 2D pose sequences into photorealistic, identity-preserving dance videos, so the key challenge is to generate temporally coherent, rhythm-aligned 2D poses from music, especially u…

Image Generation

EDGE: Editable Dance Generation From Music

2022-11-19 · CVPR 2023 1 · Jonathan Tseng, Rodrigo Castellon, C. Karen Liu

Dance is an important human art form, but creating new dances can be difficult and time-consuming. In this work, we introduce Editable Dance GEneration (EDGE), a state-of-the-art method for editable dance generation that…

DiversityMotion Synthesis

TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation

2026-04-18 · Xinran Liu, Diptesh Kanojia, Wenwu Wang, Zhenhua Feng arxiv

Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making it difficult to guide specific movements …

Cross-Modal Retrieval

MDSC: Towards Evaluating the Style Consistency Between Music and Dance

2023-09-04 · Zixiang Zhou, Weiyuan Li, Baoyuan Wang

We propose MDSC(Music-Dance-Style Consistency), the first evaluation metric that assesses to what degree the dance moves and music match. Existing metrics can only evaluate the motion fidelity and diversity and the degre…

DiversityMotion Generation