paper-with-me

홈 › Papers

LM2D: Lyrics- and Music-Driven Dance Synthesis

2024-03-14 · Wenjie Yin, Xuejiao Zhao, Yi Yu, Hang Yin, Danica Kragic, Mårten Björkman

Dance typically involves professional choreography with complex movements that follow a musical rhythm and can also be influenced by lyrical content. The integration of lyrics in addition to the auditory dimension, enriches the foundational tone and makes motion generation more amenable to its semantic meanings. However, existing dance synthesis methods tend to model motions only conditioned on audio signals. In this work, we make two contributions to bridge this gap. First, we propose LM2D, a novel probabilistic architecture that incorporates a multimodal diffusion model with consistency distillation, designed to create dance conditioned on both music and lyrics in one diffusion generation step. Second, we introduce the first 3D dance-motion dataset that encompasses both music and lyrics, obtained with pose estimation technologies. We evaluate our model against music-only baseline models with objective metrics and human evaluations, including dancers and choreographers. The results demonstrate LM2D is able to produce realistic and diverse dance matching both lyrics and music. A video summary can be accessed at: https://youtu.be/4XCgvYookvA.

📄 PDF Abstract BibTeX arXiv:2403.09407

Code (0)

등록된 구현이 없습니다.

Tasks

Motion GenerationPose EstimationRhythm

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Music- and Lyrics-driven Dance Synthesis

2023-09-30 · Wenjie Yin, Qingyuan Yao, Yi Yu, Hang Yin 외

Lyrics often convey information about the songs that are beyond the auditory dimension, enriching the semantic meaning of movements and musical themes. Such insights are important in the dance choreography domain. Howeve…

Triplet

Enabling Embodied Analogies in Intelligent Music Systems

2017-11-30 · Fabio Paolizzo

The present methodology is aimed at cross-modal machine learning and uses multidisciplinary tools and methods drawn from a broad range of areas and disciplines, including music, systematic musicology, dance, motion captu…

Audio Signal ProcessingBIG-bench Machine Learning

A Survey on Recent Deep Learning-driven Singing Voice Synthesis Systems

2021-10-06 · Yin-Ping Cho, Fu-Rong Yang, Yung-Chuan Chang, Ching-Ting Cheng 외

Singing voice synthesis (SVS) is a task that aims to generate audio signals according to musical scores and lyrics. With its multifaceted nature concerning music and language, producing singing voices indistinguishable f…

Deep LearningSinging Voice Synthesis

Lyrics Matter: Exploiting the Power of Learnt Representations for Music Popularity Prediction

2025-12-05 · Yash Choudhary, Preeti Rao, Pushpak Bhattacharyya arxiv

Accurately predicting music popularity is a critical challenge in the music industry, offering benefits to artists, producers, and streaming platforms. Prior research has largely focused on audio features, social metadat…

Music-driven Dance Regeneration with Controllable Key Pose Constraints

2022-07-08 · Junfu Pu, Ying Shan

In this paper, we propose a novel framework for music-driven dance motion synthesis with controllable key pose constraint. In contrast to methods that generate dance motion sequences only based on music without any other…

DecoderMotion Synthesis