paper-with-me

Papers

TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation

2025-04-11 · CVPR 2025 1 · Ruineng Li, Daitao Xing, Huiming Sun, Yuanzhou Ha, Jinglin Shen, Chiuman Ho

Human-centric motion control in video generation remains a critical challenge, particularly when jointly controlling camera movements and human poses in scenarios like the iconic Grammy Glambot moment. While recent video diffusion models have made significant progress, existing approaches struggle with limited motion representations and inadequate integration of camera and human motion controls. In this work, we present TokenMotion, the first DiT-based video diffusion framework that enables fine-grained control over camera motion, human motion, and their joint interaction. We represent camera trajectories and human poses as spatio-temporal tokens to enable local control granularity. Our approach introduces a unified modeling framework utilizing a decouple-and-fuse strategy, bridged by a human-aware dynamic mask that effectively handles the spatially-and-temporally varying nature of combined motion signals. Through extensive experiments, we demonstrate TokenMotion's effectiveness across both text-to-video and image-to-video paradigms, consistently outperforming current state-of-the-art methods in human-centric motion control tasks. Our work represents a significant advancement in controllable video generation, with particular relevance for creative production applications.

📄 PDF Abstract BibTeX arXiv:2504.08181

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

TokenMotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection Via Learnable Token Selection

2023-11-05 · Zifan Yu, Erfan Bank Tavakoli, Meida Chen, Suya You 외

The area of Video Camouflaged Object Detection (VCOD) presents unique challenges in the field of computer vision due to texture similarities between target objects and their surroundings, as well as irregular motion patt…

object-detectionObject Detection

Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion

2025-02-11 · Xingpei Ma, Jiaran Cai, Yuansheng Guan, Shenneng Huang 외

Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches …

AttributeDisentanglementFace GenerationPortrait Animation+1

MotionMaster: Training-free Camera Motion Transfer For Video Generation

2024-04-24 · Teng Hu, Jiangning Zhang, Ran Yi, Yating Wang 외

The emergence of diffusion models has greatly propelled the progress in image and video generation. Recently, some efforts have been made in controllable video generation, including text-to-video generation and video mot…

DisentanglementMotion DisentanglementText-to-Video GenerationVideo Generation

SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head Generation

2024-05-12 · Changpeng Cai, Guinan Guo, Jiao Li, Junhao Su 외

Most earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. Whil…

DisentanglementFace GenerationTalking Face GenerationTalking Head Generation

LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

2025-07-08 · Yisu Zhang, Chenjie Cao, Chaohui Yu, Jianke Zhu

Video Diffusion Models (VDMs) have demonstrated remarkable capabilities in synthesizing realistic videos by learning from large-scale data. Although vanilla Low-Rank Adaptation (LoRA) can learn specific spatial or tempor…