paper-with-me

Papers

MusicInfuser: Making Video Diffusion Listen and Dance

2025-03-18 · Susung Hong, Ira Kemelmacher-Shlizerman, Brian Curless, Steven M. Seitz

We introduce MusicInfuser, an approach for generating high-quality dance videos that are synchronized to a specified music track. Rather than attempting to design and train a new multimodal audio-video model, we show how existing video diffusion models can be adapted to align with musical inputs by introducing lightweight music-video cross-attention and a low-rank adapter. Unlike prior work requiring motion capture data, our approach fine-tunes only on dance videos. MusicInfuser achieves high-quality music-driven video generation while preserving the flexibility and generative capabilities of the underlying models. We introduce an evaluation framework using Video-LLMs to assess multiple dimensions of dance generation quality. The project page and code are available at https://susunghong.github.io/MusicInfuser.

📄 PDF Abstract BibTeX arXiv:2503.14505

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Every Image Listens, Every Image Dances: Music-Driven Image Animation

2025-01-30 · Zhikang Dong, Weituo Hao, Ju-Chiang Wang, Peng Zhang 외

Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video generation guided by text, music-driven d…

Image AnimationVideo Generation

Diffusion-based Realistic Listening Head Generation via Hybrid Motion Modeling

2025-01-01 · CVPR 2025 1 · Yinuo Wang, Yanbo Fan, Xuan Wang, Guo Yu 외

Listening head generation aims to synthesize non-verbal responsive listening head videos that naturally react to a certain speaker, for which, both realistic head movements, expressive facial expressions, and high vi…

Motion GenerationVideo Generation

Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models

2022-11-17 · Simon Alexanderson, Rajmund Nagy, Jonas Beskow, Gustav Eje Henter

Diffusion models have experienced a surge of interest as highly expressive yet efficiently trainable probabilistic models. We show that these models are an excellent fit for synthesising human motion that co-occurs with …

Gesture GenerationMotion Synthesis

MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion Model

2023-08-31 · Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai 외

Face-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlook…

DenoisingDiversity

SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning

2026-08-17 · Tao Feng, Xu Li, Xiangyang Luo, Ming Wen 외 arxiv

Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate th…

Video Generation