paper-with-me

홈 › Papers

ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer

2025-04-03 · CVPR 2025 1 · Jiayi Gao, Zijin Yin, Changcheng Hua, Yuxin Peng, Kongming Liang, Zhanyu Ma, Jun Guo, Yang Liu

The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to preserve the diversity and accuracy of motion as transferring to subjects with varying shapes. To overcome these, we introduce \textbf{ConMo}, a zero-shot framework that disentangle and recompose the motions of subjects and camera movements. ConMo isolates individual subject and background motion cues from complex trajectories in source videos using only subject masks, and reassembles them for target video generation. This approach enables more accurate motion control across diverse subjects and improves performance in multi-subject scenarios. Additionally, we propose soft guidance in the recomposition stage which controls the retention of original motion to adjust shape constraints, aiding subject shape adaptation and semantic transformation. Unlike previous methods, ConMo unlocks a wide range of applications, including subject size and position editing, subject removal, semantic modifications, and camera motion simulation. Extensive experiments demonstrate that ConMo significantly outperforms state-of-the-art methods in motion fidelity and semantic consistency. The code is available at https://github.com/Andyplus1/ConMo.

📄 PDF Abstract BibTeX arXiv:2504.02451

Code (1)

andyplus1/conmo 공식 구현 pytorch

Tasks

DisentanglementMotion DisentanglementVideo Generation

Similar Papers 제목 키워드 기반

VAW-GAN for Disentanglement and Recomposition of Emotional Elements in Speech

2020-11-03 · Kun Zhou, Berrak Sisman, Haizhou Li

Emotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition…

DecoderDisentanglementGenerative Adversarial NetworkVoice Conversion

CONMOD: Controllable Neural Frame-based Modulation Effects

2024-06-20 · Gyubin Lee, Hounsu Kim, Junwon Lee, Juhan Nam

Deep learning models have seen widespread use in modelling LFO-driven audio effects, such as phaser and flanger. Although existing neural architectures exhibit high-quality emulation of individual effects, they do not po…

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment

2026-04-21 · Chaonan Ji, Jinwei Qi, Sheng Xu, Peng Zhang 외 arxiv

Existing facial reenactment methods struggle with a trade-off between expressiveness and fine-grained controllability. Holistic facial reenactment models often sacrifice granular control for expressiveness, while methods…

EMOSH: Expressive Motion and Shape Disentanglement for Human Animation

2026-06-26 · Dongbin Zhang, Hao Liu, Binquan Dai, Kangjie Chen 외 arxiv

High-fidelity and expressive controllable human animation is essential for content creation and digital avatar applications. However, existing methods face a dilemma between expressiveness and disentanglement. Mainstream…

Video Generation

Encouraging Disentangled and Convex Representation with Controllable Interpolation Regularization

2021-12-06 · Yunhao Ge, Zhi Xu, Yao Xiao, Gan Xin 외

We focus on controllable disentangled representation learning (C-Dis-RL), where users can control the partition of the disentangled latent space to factorize dataset attributes (concepts) for downstream tasks. Two genera…

Data AugmentationDisentanglementFairnessImage Generation+1