paper-with-me

홈 › Papers

From Diffusion to Flow: Efficient Motion Generation in MotionGPT3

2026-03-23 · Jaymin Ban, JiHong Jeon, SangYeop Jeong arxiv

Recent text-driven motion generation methods span both discrete token-based approaches and continuous-latent formulations. MotionGPT3 exemplifies the latter paradigm, combining a learned continuous motion latent space with a diffusion-based prior for text-conditioned synthesis. While rectified flow objectives have recently demonstrated favorable convergence and inference-time properties relative to diffusion in image and audio generation, it remains unclear whether these advantages transfer cleanly to the motion generation setting. In this work, we conduct a controlled empirical study comparing diffusion and rectified flow objectives within the MotionGPT3 framework. By holding the model architecture, training protocol, and evaluation setup fixed, we isolate the effect of the generative objective on training dynamics, final performance, and inference efficiency. Experiments on the HumanML3D dataset show that rectified flow converges in fewer training epochs, reaches strong test performance earlier, and matches or exceeds diffusion-based motion quality under identical conditions. Moreover, flow-based priors exhibit stable behavior across a wide range of inference step counts and achieve competitive quality with fewer sampling steps, yielding improved efficiency-quality trade-offs. Overall, our results suggest that several known benefits of rectified flow objectives do extend to continuous-latent text-to-motion generation, highlighting the importance of the training objective choice in motion priors.

📄 PDF Abstract BibTeX arXiv:2603.26747

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Similar Papers 제목 키워드 기반

MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

2024-10-29 · YuAn Wang, Di Huang, Yaqi Zhang, Wanli Ouyang 외

Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Despite impressive advances, existing appro…

DescriptiveLanguage ModelingLanguage ModellingMotion Captioning+2

MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators

2023-06-19 · Yaqi Zhang, Di Huang, Bin Liu, Shixiang Tang 외

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in gene…

Motion Generation

MotionGPT: Human Motion as a Foreign Language

2023-06-26 · NeurIPS 2023 11 · Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu 외

Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains challenging and untouched so far. Fortunat…

Language ModelingLanguage ModellingMotion CaptioningMotion Generation+3

Exploring Text-to-Motion Generation with Human Preference

2024-04-15 · Jenny Sheng, Matthieu Lin, Andrew Zhao, Kevin Pruvost 외

This paper presents an exploration of preference learning in text-to-motion generation. We find that current improvements in text-to-motion generation still rely on datasets requiring expert labelers with motion capture …

Motion Generation

OmniMotionGPT: Animal Motion Generation with Limited Data

2023-11-30 · CVPR 2024 1 · Zhangsihao Yang, Mingyuan Zhou, Mengyi Shan, Bingbing Wen 외

Our paper aims to generate diverse and realistic animal motion sequences from textual descriptions, without a large-scale animal text-motion dataset. While the task of text-driven human motion synthesis is already extens…

DiversityMotion GenerationMotion Synthesis