paper-with-me

Papers

OmniDance: Multimodal Driven Dance Video Generation with Large-scale Internet Data

2026-06-29 · Kaixing Yang, Jiashu Zhu, Xulong Tang, Ziqiao Peng, Xiangyue Zhang, Chubin Chen, Puwei Wang, Jiahong Wu, Xiangxiang Chu, Hongyan Liu, Jun He arxiv

Music-driven dance video generation aims to synthesize expressive human motion that is temporally aligned with music while maintaining high visual fidelity. Despite recent progress, existing methods still face two key limitations: the lack of large-scale, high-quality dance video datasets, and the absence of principled frameworks for integrating music as a complementary conditioning signal into Video Generation Foundation Models. To address these limitations, we introduce CIPE-Dance, a large-scale Internet-sourced dance video dataset with choreography-informed text annotations, constructed via a progressive expert pipeline. To the best of our knowledge, CIPE-Dance is the largest dataset for dance video generation to date, comprising 300k high-quality clips over 400 hours and covering diverse dancers, environments, and dance genres. We further propose OmniDance, a framework-level recipe for integrating music into a TI2V foundation model without sacrificing its original controllability or visual fidelity. Motivated by the complementary roles of text as low-frequency semantics and music as high-frequency temporal dynamics, OmniDance co-designs a depth-aware specialization architecture, an anchored easy-to-hard curriculum learning strategy, and a modality-specialized time-dependent CFG strategy, enabling unified TI2V, MI2V, and MTI2V generation. Extensive experiments on CIPE-Dance demonstrate that OmniDance achieves state-of-the-art performance across all three tasks and exhibits robust multimodal integration capability. Project is available at https://github.com/AMAP-ML/OmniDance.

📄 PDF Abstract BibTeX arXiv:2606.30019

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Every Image Listens, Every Image Dances: Music-Driven Image Animation

2025-01-30 · Zhikang Dong, Weituo Hao, Ju-Chiang Wang, Peng Zhang 외

Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video generation guided by text, music-driven d…

Image AnimationVideo Generation

MusicInfuser: Making Video Diffusion Listen and Dance

2025-03-18 · Susung Hong, Ira Kemelmacher-Shlizerman, Brian Curless, Steven M. Seitz

We introduce MusicInfuser, an approach for generating high-quality dance videos that are synchronized to a specified music track. Rather than attempting to design and train a new multimodal audio-video model, we show how…

Video Generation

OpenDance: Multimodal Controllable 3D Dance Generation Using Large-scale Internet Data

2025-06-09 · Jinlu Zhang, Zixi Kang, Yizhou Wang

Music-driven dance generation offers significant creative potential yet faces considerable challenges. The absence of fine-grained multimodal data and the difficulty of flexible multi-conditional generation limit previou…

Diversity

Stance-Driven Multimodal Controlled Statement Generation: New Dataset and Task

2025-04-04 · Bingqian Wang, Quan Fang, Jiachen Sun, Xiaoxiao Ma

Formulating statements that support diverse or controversial stances on specific topics is vital for platforms that enable user expression, reshape political discourse, and drive social critique and information dissemina…

Marketingmultimodal generationStance DetectionText Generation

HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation

2025-06-10 · Ziyao Huang, Zixiang Zhou, Juan Cao, Yifeng Ma 외

To address key limitations in human-object interaction (HOI) video generation -- specifically the reliance on curated motion data, limited generalization to novel objects/scenarios, and restricted accessibility -- we int…

Human AnimationHuman-Object Interaction DetectionVideo Generation