MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
Multi-object video motion transfer poses significant challenges for Diffusion Transformer (DiT) architectures due to inherent motion entanglement and lack of object-level control. We present MultiMotion, a novel unified framework that overcomes these limitations. Our core innovation is Maskaware Attention Motion Flow (AMF), which utilizes SAM2 masks to explicitly disentangle and control motion features for multiple objects within the DiT pipeline. Furthermore, we introduce RectPC, a high-order predictor-corrector solver for efficient and accurate sampling, particularly beneficial for multi-entity generation. To facilitate rigorous evaluation, we construct the first benchmark dataset specifically for DiT-based multi-object motion transfer. MultiMotion demonstrably achieves precise, semantically aligned, and temporally coherent motion transfer for multiple distinct objects, maintaining DiT's high quality and scalability. The code is in the supp.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle mu…
DisentanglementMotion DisentanglementVideo Generation3D-Aware Talking-Head Video Motion Transfer
Motion transfer of talking-head videos involves generating a new video with the appearance of a subject video and the motion pattern of a driving video. Current methodologies primarily depend on a limited number of subje…
Novel View SynthesisEverybody Dance Now
This paper presents a simple method for "do as I do" motion transfer: given a source video of a person dancing, we can transfer that performance to a novel (amateur) target after only a few minutes of the target subject …
Face GenerationImage-to-Image TranslationVideo GenerationFlow Guided Transformable Bottleneck Networks for Motion Retargeting
Human motion retargeting aims to transfer the motion of one person in a "driving" video or set of images to another person. Existing efforts leverage a long training video from each target person to train a subject-speci…
Image Generationmotion retargetingClusterVO: Clustering Moving Instances and Estimating Visual Odometry for Self and Surroundings
We present ClusterVO, a stereo Visual Odometry which simultaneously clusters and estimates the motion of both ego and surrounding rigid clusters/objects. Unlike previous solutions relying on batch input or imposing prior…
Autonomous DrivingClusteringScene UnderstandingTrajectory Recovery+1