paper-with-me

Papers

TeMuDance: Contrastive Alignment-Based Textual Control for Music-Driven Dance Generation

2026-04-18 · Xinran Liu, Diptesh Kanojia, Wenwu Wang, Zhenhua Feng arxiv

Existing music-driven dance generation approaches have achieved strong realism and effective audio-motion alignment. However, they generally lack semantic controllability, making it difficult to guide specific movements through natural language descriptions. This limitation primarily stems from the absence of large-scale datasets that jointly align music, text, and motion for supervised learning of text-conditioned control. To address this challenge, we propose TeMuDance, a framework that enables text-based control for music-conditioned dance generation without requiring any manually annotated music-text-motion triplet dataset. TeMuDance introduces a motion-centred bridging paradigm that leverages motion as a shared semantic anchor to align disjoint music-dance and text-motion datasets within a unified embedding space, enabling cross-modal retrieval of missing modalities for end-to-end training. A lightweight text control branch is then trained on top of a frozen music-to-dance diffusion backbone, preserving rhythmic fidelity while enabling fine-grained semantic guidance. To further suppress noise inherent in the retrieved supervision, we design a dual-stream fine-tuning strategy with confidence-based filtering. We also propose a novel task-aligned metric that quantifies whether textual prompts induce the intended kinematic attributes under music conditioning. Extensive experiments demonstrate that TeMuDance achieves competitive dance quality while substantially improving text-conditioned control over existing methods.

📄 PDF Abstract BibTeX arXiv:2604.17005

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal Retrieval

Similar Papers 제목 키워드 기반

MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

2024-10-16 · RuiQi Li, Siqi Zheng, Xize Cheng, Ziang Zhang 외

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmon…

In-Context LearningMusic GenerationRhythm

Contrastive Audio-Language Learning for Music

2022-08-25 · Ilaria Manco, Emmanouil Benetos, Elio Quinton, György Fazekas

As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information R…

Audio to Text RetrievalDescriptiveGenre classificationInformation Retrieval+3

Multimodal Music Generation with Explicit Bridges and Retrieval Augmentation

2024-12-12 · Baisen Wang, Le Zhuo, Zhaokai Wang, Chenxi Bao 외

Multimodal music generation aims to produce music from diverse input modalities, including text, videos, and images. Existing methods use a common embedding space for multimodal fusion. Despite their effectiveness in oth…

cross-modal alignmentMultimodal Music GenerationMusic GenerationRetrieval

Controllable Video-to-Music Generation with Multiple Time-Varying Conditions

2025-07-28 · Junxian Wu, Weitao You, Heda Zuo, Dengming Zhang 외 arxiv

Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music i…

Music Generation

EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation

2025-07-07 · Fathinah Izzati, Xinyue Li, Gus Xia arxiv

We propose Expotion (Facial Expression and Motion Control for Multimodal Music Generation), a generative model leveraging multimodal visual controls - specifically, human facial expressions and upper-body motion - as wel…

parameter-efficient fine-tuningText-to-Music Generation