paper-with-me

홈 › Papers

MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer

2026-04-01 · Samuel Teodoro, Yun Chen, Agus Gunawan, Soo Ye Kim, Jihyong Oh, Munchurl Kim arxiv

Motion transfer enables controllable video generation by transferring temporal dynamics from a reference video to synthesize a new video conditioned on a target caption. However, existing Diffusion Transformer (DiT)-based methods are limited to single-object videos, restricting fine-grained control in real-world scenes with multiple objects. In this work, we introduce MotionGrounder, a DiT-based framework that firstly handles motion transfer with multi-object controllability. Our Flow-based Motion Signal (FMS) in MotionGrounder provides a stable motion prior for target video generation, while our Object-Caption Alignment Loss (OCAL) grounds object captions to their corresponding spatial regions. We further propose a new Object Grounding Score (OGS), which jointly evaluates (i) spatial alignment between source video objects and their generated counterparts and (ii) semantic consistency between each generated object and its target caption. Our experiments show that MotionGrounder consistently outperforms recent baselines across quantitative, qualitative, and human evaluations.

📄 PDF Abstract BibTeX arXiv:2604.00853

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

MotionShop: Zero-Shot Motion Transfer in Video Diffusion Models with Mixture of Score Guidance

2024-12-06 · Hidir Yesiltepe, Tuna Han Salih Meral, Connor Dunlop, Pinar Yanardag

In this work, we propose the first motion transfer approach in diffusion transformer through Mixture of Score Guidance (MSG), a theoretically-grounded framework for motion transfer in diffusion models. Our key theoretica…

Object

GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction

2026-06-15 · Ruoxuan Yang, Tieyuan Chen, Xiaofeng Huang, Haibing Yin 외 arxiv

Viewer sentiment prediction in video advertisements aims to infer the latent affective response evoked in the audience. To bridge the gap between what is shown and what is felt, models must deduce hidden viewer emotions …

HACMan++: Spatially-Grounded Motion Primitives for Manipulation

2024-07-11 · Bowen Jiang, Yilin Wu, Wenxuan Zhou, Chris Paxton 외

Although end-to-end robot learning has shown some success for robot manipulation, the learned policies are often not sufficiently robust to variations in object pose or geometry. To improve the policy generalization, we …

ObjectRobot Manipulation

Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer

2026-03-01 · Yuze Li, Dong Gong, Xiao Cao, Junchao Yuan 외 arxiv

Motion transfer has emerged as a promising direction for controllable video generation, yet existing methods largely focus on single-object scenarios and struggle when multiple objects require distinct motion patterns. I…

Video Generation

MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer

2025-12-08 · Penghui Liu, Jiangshan Wang, Yutong Shen, Shanhui Mo 외 arxiv

Multi-object video motion transfer poses significant challenges for Diffusion Transformer (DiT) architectures due to inherent motion entanglement and lack of object-level control. We present MultiMotion, a novel unified …