paper-with-me

Papers

X-MoGen: Unified Motion Generation across Humans and Animals

2025-08-07 · Xuan Wang, Kai Ruan, Liyang Qian, Zhizhi Guo, Chang Su, Gaoang Wang arxiv

Text-driven motion generation has attracted increasing attention due to its broad applications in virtual reality, animation, and robotics. While existing methods typically model human and animal motion separately, a joint cross-species approach offers key advantages, such as a unified representation and improved generalization. However, morphological differences across species remain a key challenge, often compromising motion plausibility. To address this, we propose X-MoGen, the first unified framework for cross-species text-driven motion generation covering both humans and animals. X-MoGen adopts a two-stage architecture. First, a conditional graph variational autoencoder learns canonical T-pose priors, while an autoencoder encodes motion into a shared latent space regularized by morphological loss. In the second stage, we perform masked motion modeling to generate motion embeddings conditioned on textual descriptions. During training, a morphological consistency module is employed to promote skeletal plausibility across species. To support unified modeling, we construct UniMo4D, a large-scale dataset of 115 species and 119k motion sequences, which integrates human and animal motions under a shared skeletal topology for joint training. Extensive experiments on UniMo4D demonstrate that X-MoGen outperforms state-of-the-art methods on both seen and unseen species.

📄 PDF Abstract BibTeX arXiv:2508.05162

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions

2025-12-22 · Wendong Bu, Kaihang Pan, Yuze Lin, Jiacheng Li 외 arxiv

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting …

CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration

2026-05-21 · Adil Meric, Lin Geng Foo, Mert Kiray, Benjamin Busam 외 arxiv

We present CoMoGen, a controllable video generation framework that generates realistic interactive dynamics from a single binary mask sequence conditioned on an input image. CoMoGen introduces a lightweight MaskAdapter t…

Video Generation

The Quest for Generalizable Motion Generation: Data, Model, and Evaluation

2025-10-30 · Jing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang 외 arxiv

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generativ…

Video Generation

ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation

2026-05-12 · Inwoo Hwang, Hojun Jang, Bing Zhou, Jian Wang 외 arxiv

We present ScaleMoGen, a scale-wise autoregressive framework for text-driven human motion generation. Unlike conventional autoregressive approaches that rely on standard next-token prediction, ScaleMoGen frames motion ge…

InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation

2026-07-11 · Yichen Peng, Jyun-Ting Song, Chen-Chieh Liao, Kris Kitani 외 arxiv

Human-pet interaction estimation and generation remain underexplored due to the absence of a high-quality large-scale dataset. We present InterPet4D, the first multimodal dataset capturing natural interactions between hu…