paper-with-me

홈 › Papers

HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions

2025-05-29 · Shuolin Xu, Siming Zheng, Ziyi Wang, HC Yu, Jinwei Chen, Huaqi Zhang, Bo Li, Peng-Tao Jiang

Recent advances in diffusion models have significantly improved conditional video generation, particularly in the pose-guided human image animation task. Although existing methods are capable of generating high-fidelity and time-consistent animation sequences in regular motions and static scenes, there are still obvious limitations when facing complex human body motions (Hypermotion) that contain highly dynamic, non-standard motions, and the lack of a high-quality benchmark for evaluation of complex human motion animations. To address this challenge, we introduce the \textbf{Open-HyperMotionX Dataset} and \textbf{HyperMotionX Bench}, which provide high-quality human pose annotations and curated video clips for evaluating and improving pose-guided human image animation models under complex human motion conditions. Furthermore, we propose a simple yet powerful DiT-based video generation baseline and design spatial low-frequency enhanced RoPE, a novel module that selectively enhances low-frequency spatial feature modeling by introducing learnable frequency scaling. Our method significantly improves structural stability and appearance consistency in highly dynamic human motion sequences. Extensive experiments demonstrate the effectiveness of our dataset and proposed approach in advancing the generation quality of complex human motion image animations. Code and dataset will be made publicly available.

📄 PDF Abstract BibTeX arXiv:2505.22977

Code (1)

vivoCameraResearch/Hyper-Motion 공식 구현 pytorch

Tasks

Image AnimationVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MagicAvatar: Multimodal Avatar Generation and Animation

2023-08-28 · Jianfeng Zhang, Hanshu Yan, Zhongcong Xu, Jiashi Feng 외

This report presents MagicAvatar, a framework for multimodal video generation and animation of human avatars. Unlike most existing methods that generate avatar-centric videos directly from multimodal inputs (e.g., text p…

Video Generation

QS-Craft: Learning to Quantize, Scrabble and Craft for Conditional Human Motion Animation

2022-03-22 · Yuxin Hong, Xuelin Qian, Simian Luo, xiangyang xue 외

This paper studies the task of conditional Human Motion Animation (cHMA). Given a source image and a driving video, the model should animate the new frame sequence, in which the person in the source image should perform …

Generative Adversarial Network

MultiAnimate: Pose-Guided Image Animation Made Extensible

2026-02-25 · Yingcheng Hu, Haowen Gong, Chuanguang Yang, Zhulin An 외 arxiv

Pose-guided human image animation aims to synthesize realistic videos of a reference character driven by a sequence of poses. While diffusion-based methods have achieved remarkable success, most existing approaches are l…

Video Generation

MTVCrafter: 4D Motion Tokenization for Open-World Human Image Animation

2025-05-15 · Yanbo Ding, Xirui Hu, Zhizhi Guo, Yali Wang

Human image animation has gained increasing attention and developed rapidly due to its broad applications in digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which li…

Image AnimationVideo Generation

Make-An-Animation: Large-Scale Text-conditional 3D Human Motion Generation

2023-05-16 · ICCV 2023 1 · Samaneh Azadi, Akbar Shah, Thomas Hayes, Devi Parikh 외

Text-guided human motion generation has drawn significant interest because of its impactful applications spanning animation and robotics. Recently, application of diffusion models for motion generation has enabled improv…

Motion GenerationMotion SynthesisText-to-Video GenerationVideo Generation