paper-with-me

Papers

Every Image Listens, Every Image Dances: Music-Driven Image Animation

2025-01-30 · Zhikang Dong, Weituo Hao, Ju-Chiang Wang, Peng Zhang, Pawel Polak

Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video generation guided by text, music-driven dance video generation remains underexplored. In this paper, we introduce MuseDance, an innovative end-to-end model that animates reference images using both music and text inputs. This dual input enables MuseDance to generate personalized videos that follow text descriptions and synchronize character movements with the music. Unlike existing approaches, MuseDance eliminates the need for complex motion guidance inputs, such as pose or depth sequences, making flexible and creative video generation accessible to users of all expertise levels. To advance research in this field, we present a new multimodal dataset comprising 2,904 dance videos with corresponding background music and text descriptions. Our approach leverages diffusion-based methods to achieve robust generalization, precise control, and temporal consistency, setting a new baseline for the music-driven image animation task.

📄 PDF Abstract BibTeX arXiv:2501.18801

Code (0)

등록된 구현이 없습니다.

Tasks

Image AnimationVideo Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Cooperative Graph Neural Networks

2023-10-02 · Ben Finkelshtein, Xingyue Huang, Michael Bronstein, İsmail İlkan Ceylan

Graph neural networks are popular architectures for graph machine learning, based on iterative computation of node representations of an input graph through a series of invariant transformations. A large class of graph n…

Augmenting Human Cognition through Everyday AR

2025-05-06 · Xiaoan Liu

As spatial computing and multimodal LLMs mature, AR is tending to become an intuitive "thinking tool," embedding semantic and context-aware intelligence directly into everyday environments. This paper explores how always…

GrASP: Gradient-Based Affordance Selection for Planning

2022-02-08 · Vivek Veeriah, Zeyu Zheng, Richard Lewis, Satinder Singh

Planning with a learned model is arguably a key component of intelligence. There are several challenges in realizing such a component in large-scale reinforcement learning (RL) problems. One such challenge is dealing eff…

Reinforcement Learning (RL)

Edit Everything: A Text-Guided Generative System for Images Editing

2023-04-27 · Defeng Xie, Ruichen Wang, Jian Ma, Chen Chen 외

We introduce a new generative system called Edit Everything, which can take image and text inputs and produce image outputs. Edit Everything allows users to edit images using simple text instructions. Our system designs …

Leveraging Visual Knowledge in Language Tasks: An Empirical Study on Intermediate Pre-training for Cross-modal Knowledge Transfer

2022-03-14 · ACL 2022 5 · Woojeong Jin, Dong-Ho Lee, Chenguang Zhu, Jay Pujara 외

Pre-trained language models are still far from human performance in tasks that need understanding of properties (e.g. appearance, measurable quantity) and affordances of everyday objects in the real world since the text …

Image CaptioningLanguage ModelingLanguage ModellingTransfer Learning