paper-with-me

홈 › Papers

CoMo: Compositional Motion Customization for Text-to-Video Generation

2025-10-27 · Youcan Xu, Zhen Wang, Jiaxin Shi, Kexin Li, Feifei Shao, Jun Xiao, Yi Yang, Jun Yu, Long Chen arxiv

While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods for single-motion customization have been developed to address this gap, they fail in compositional scenarios due to two primary challenges: motion-appearance entanglement and ineffective multi-motion blending. This paper introduces CoMo, a novel framework for $\textbf{compositional motion customization}$ in text-to-video generation, enabling the synthesis of multiple, distinct motions within a single video. CoMo addresses these issues through a two-phase approach. First, in the single-motion learning phase, a static-dynamic decoupled tuning paradigm disentangles motion from appearance to learn a motion-specific module. Second, in the multi-motion composition phase, a plug-and-play divide-and-merge strategy composes these learned motions without additional training by spatially isolating their influence during the denoising process. To facilitate research in this new domain, we also introduce a new benchmark and a novel evaluation metric designed to assess multi-motion fidelity and blending. Extensive experiments demonstrate that CoMo achieves state-of-the-art performance, significantly advancing the capabilities of controllable video generation. Our project page is at https://como6.github.io/.

📄 PDF Abstract BibTeX arXiv:2510.23007

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models

2024-02-22 · Yixuan Ren, Yang Zhou, Jimei Yang, Jing Shi 외

Image customization has been extensively studied in text-to-image (T2I) diffusion models, leading to impressive outcomes and applications. With the emergence of text-to-video (T2V) diffusion models, its temporal counterp…

Video Generation

VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion Models

2023-12-01 · CVPR 2024 1 · Hyeonho Jeong, Geon Yeong Park, Jong Chul Ye

Text-to-video diffusion models have advanced video generation significantly. However, customizing these models to generate videos with tailored motions presents a substantial challenge. In specific, they encounter hurdle…

Video EditingVideo Generation

MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching

2025-02-18 · Yen-Siang Wu, Chi-Pin Huang, Fu-En Yang, Yu-Chiang Frank Wang

Text-to-video (T2V) diffusion models have shown promising capabilities in synthesizing realistic videos from input text prompts. However, the input text description alone provides limited control over the precise objects…

Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization

2026-06-25 · Xuancheng Xu, Gengyun Jia, Bing-Kun Bao arxiv

Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos. While significant strides have been made in image stylization and video motion …

Text-to-Video Generation

NewMove: Customizing text-to-video models with novel motions

2023-12-07 · Joanna Materzynska, Josef Sivic, Eli Shechtman, Antonio Torralba 외

We introduce an approach for augmenting text-to-video generation models with customized motions, extending their capabilities beyond the motions depicted in the original training data. By leveraging a few video samples d…

Text-to-Video GenerationVideo Generation