paper-with-me

홈 › Papers

MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm

2025-02-04 · Ziyan Guo, Zeyu Hu, Na Zhao, De Wen Soh

Human motion generation and editing are key components of computer graphics and vision. However, current approaches in this field tend to offer isolated solutions tailored to specific tasks, which can be inefficient and impractical for real-world applications. While some efforts have aimed to unify motion-related tasks, these methods simply use different modalities as conditions to guide motion generation. Consequently, they lack editing capabilities, fine-grained control, and fail to facilitate knowledge sharing across tasks. To address these limitations and provide a versatile, unified framework capable of handling both human motion generation and editing, we introduce a novel paradigm: Motion-Condition-Motion, which enables the unified formulation of diverse tasks with three concepts: source motion, condition, and target motion.Based on this paradigm, we propose a unified framework, MotionLab, which incorporates rectified flows to learn the mapping from source motion to target motion, guided by the specified conditions.In MotionLab, we introduce the 1) MotionFlow Transformer to enhance conditional generation and editing without task-specific modules; 2) Aligned Rotational Position Encoding} to guarantee the time synchronization between source motion and target motion; 3) Task Specified Instruction Modulation; and 4) Motion Curriculum Learning for effective multi-task learning and knowledge sharing across tasks. Notably, our MotionLab demonstrates promising generalization capabilities and inference efficiency across multiple benchmarks for human motion. Our code and additional video results are available at: https://diouo.github.io/motionlab.github.io/.

📄 PDF Abstract BibTeX arXiv:2502.02358

Code (0)

등록된 구현이 없습니다.

Tasks

Motion GenerationMulti-Task Learning

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Residual Connection 설명 없음
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding

2026-02-04 · Ning Zhang, Zhengyu Li, Kwong Weng Loh, Mingxi Xu 외 arxiv

Prior masked modeling motion generation methods predominantly study text-to-motion. We present DiMo, a discrete diffusion-style framework, which extends masked modeling to bidirectional text--motion understanding and gen…

NECromancer: Breathing Life into Skeletons via BVH Animation

2026-02-06 · Mingxi Xu, Qi Wang, Zhengyu Wen, Phong Dao Thien 외 arxiv

Motion tokenization is a key component of generalizable motion models, yet most existing approaches are restricted to species-specific skeletons, limiting their applicability across diverse morphologies. We propose NECro…

OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions

2025-12-22 · Wendong Bu, Kaihang Pan, Yuze Lin, Jiacheng Li 외 arxiv

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting …

MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos

2025-12-11 · Kehong Gong, Zhengyu Wen, Weixia He, Mingxi Xu 외 arxiv

Motion capture now underpins content creation far beyond digital humans, yet most existing pipelines remain species- or template-specific. We formalize this gap as Category-Agnostic Motion Capture (CAMoCap): given a mono…

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

2026-05-18 · Bizhu Wu, Jinheng Xie, Wenting Chen, Zhe Kong 외 arxiv

Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and nuanced control over body parts needed for animation or interaction.…

Zero-shot Generalization