paper-with-me

Papers

MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls

2024-07-30 · Yuxuan Bian, Ailing Zeng, Xuan Ju, Xian Liu, Zhaoyang Zhang, Wei Liu, Qiang Xu

Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to achieve various generation tasks with different condition modalities presents two main challenges: motion distribution drifts across different tasks (e.g., co-speech gestures and text-driven daily actions) and the complex optimization of mixed conditions with varying granularities (e.g., text and audio). Additionally, inconsistent motion formats across different tasks and datasets hinder effective training toward multimodal motion generation. In this paper, we propose MotionCraft, a unified diffusion transformer that crafts whole-body motion with plug-and-play multimodal control. Our framework employs a coarse-to-fine training strategy, starting with the first stage of text-to-motion semantic pre-training, followed by the second stage of multimodal low-level control adaptation to handle conditions of varying granularities. To effectively learn and transfer motion knowledge across different distributions, we design MC-Attn for parallel modeling of static and dynamic human topology graphs. To overcome the motion format inconsistency of existing benchmarks, we introduce MC-Bench, the first available multimodal whole-body motion generation benchmark based on the unified SMPL-X format. Extensive experiments show that MotionCraft achieves state-of-the-art performance on various standard motion generation tasks.

📄 PDF Abstract BibTeX arXiv:2407.21136

Code (1)

cure-lab/MotionCraft 공식 구현 pytorch

Tasks

Gesture GenerationMotion GenerationMotion Synthesismultimodal generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MotionCrafter: One-Shot Motion Customization of Diffusion Models

2023-12-08 · Yuxin Zhang, Fan Tang, Nisha Huang, Haibin Huang 외

The essence of a video lies in its dynamic motions, including character actions, object movements, and camera movements. While text-to-video generative diffusion models have recently advanced in creating diverse contents…

DisentanglementMotion DisentanglementText-to-Video GenerationVideo Editing+1

MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAE

2026-02-09 · Ruijie Zhu, Jiahao Lu, Wenbo Hu, Xiaoguang Han 외 arxiv

We present MotionCrafter, a framework that leverages video generators to jointly reconstruct 4D geometry and estimate dense motion from a monocular video. The key idea is a joint representation of dense 3D point maps and…

Scene Flow Estimation

MotionCraft: Physics-based Zero-Shot Video Generation

2024-05-22 · Luca Savant Aira, Antonio Montanaro, Emanuele Aiello, Diego Valsesia 외

Generating videos with realistic and physically plausible motion is one of the main recent challenges in computer vision. While diffusion models are achieving compelling results in image generation, video diffusion model…

Image GenerationMissing ElementsOptical Flow EstimationVideo Generation

COOP: Decoupling and Coupling of Whole-Body Grasping Pose Generation

2023-01-01 · ICCV 2023 1 · Yanzhao Zheng, Yunzhou Shi, Yuhao Cui, Zhongzhou Zhao 외

Generating life-like whole-body human grasping has garnered significant attention in the field of computer graphics. Existing works have demonstrated the effectiveness of keyframe-guided motion generation framework, …

Motion Generation

Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset

2025-01-09 · Yuhong Zhang, Jing Lin, Ailing Zeng, Guanlin Wu 외

In this paper, we introduce Motion-X++, a large-scale multimodal 3D expressive whole-body human motion dataset. Existing motion datasets predominantly capture body-only poses, lacking facial expressions, hand gestures, a…

Human Mesh RecoveryMotion Generation