paper-with-me

홈 › Papers

MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

2024-10-29 · YuAn Wang, Di Huang, Yaqi Zhang, Wanli Ouyang, Jile Jiao, Xuetao Feng, Yan Zhou, Pengfei Wan, Shixiang Tang, Dan Xu

Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Despite impressive advances, existing approaches are often constrained by limited control modalities, task specificity, and focus solely on body motion representations.In this paper, we present MotionGPT-2, a unified Large Motion-Language Model (LMLM) that addresses these limitations. MotionGPT-2 accommodates multiple motion-relevant tasks and supporting multimodal control conditions through pre-trained Large Language Models (LLMs). It quantizes multimodal inputs-such as text and single-frame poses-into discrete, LLM-interpretable tokens, seamlessly integrating them into the LLM's vocabulary. These tokens are then organized into unified prompts, guiding the LLM to generate motion outputs through a pretraining-then-finetuning paradigm. We also show that the proposed MotionGPT-2 is highly adaptable to the challenging 3D holistic motion generation task, enabled by the innovative motion discretization framework, Part-Aware VQVAE, which ensures fine-grained representations of body and hand movements. Extensive experiments and visualizations validate the effectiveness of our method, demonstrating the adaptability of MotionGPT-2 across motion generation, motion captioning, and generalized motion completion tasks.

📄 PDF Abstract BibTeX arXiv:2410.21747

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveLanguage ModelingLanguage ModellingMotion CaptioningMotion GenerationSpecificity

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators

2023-06-19 · Yaqi Zhang, Di Huang, Bin Liu, Shixiang Tang 외

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in gene…

Motion Generation

MotionGPT: Human Motion as a Foreign Language

2023-06-26 · NeurIPS 2023 11 · Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu 외

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains challenging and untouched so far. Fortunat…

Language ModelingLanguage ModellingMotion CaptioningMotion Generation+3

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

2026-03-23 · Jaymin Ban, JiHong Jeon, SangYeop Jeong arxiv

Recent text-driven motion generation methods span both discrete token-based approaches and continuous-latent formulations. MotionGPT3 exemplifies the latter paradigm, combining a learned continuous motion latent space wi…

Audio Generation

LingoMotion: An Interpretable and Unambiguous Symbolic Representation for Human Motion

2026-03-13 · Yao Zhang, Zhuchenyang Liu, Yu Xiao arxiv

Existing representations for human motion, such as MotionGPT, often operate as black-box latent vectors with limited interpretability and build on joint positions which can cause ambiguity. Inspired by the hierarchical s…

Generative AI-Driven High-Fidelity Human Motion Simulation

2025-07-18 · Hari Iyer, Neel Macwan, Atharva Jitendra Hude, Heejin Jeong 외 arxiv

Human motion simulation (HMS) supports cost-effective evaluation of worker behavior, safety, and productivity in industrial tasks. However, existing methods often suffer from low motion fidelity. This study introduces Ge…