paper-with-me

Papers

MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing

2025-09-28 · Ruibing Hou, Mingshuang Luo, Hongyu Pan, Hong Chang, Shiguang Shan arxiv

This paper proposes MotionVerse, a unified framework that harnesses the capabilities of Large Language Models (LLMs) to comprehend, generate, and edit human motion in both single-person and multi-person scenarios. To efficiently represent motion data, we employ a motion tokenizer with residual quantization, which converts continuous motion sequences into multi-stream discrete tokens. Furthermore, we introduce a \textit{Delay Parallel} Modeling strategy, which temporally staggers the encoding of residual token streams. This design enables LLMs to effectively capture inter-stream dependencies while maintaining computational efficiency comparable to single-stream modeling. Moreover, to alleviate modality interference between motion and language, we design a \textit{dual-tower architecture} with modality-specific parameters, ensuring stable integration of motion information for both comprehension and generation tasks. Comprehensive ablation studies demonstrate the effectiveness of each component in MotionVerse, and extensive experiments showcase its superior performance across a wide range of motion-relevant tasks.

📄 PDF Abstract BibTeX arXiv:2509.23635

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

MotionLLaMA: A Unified Framework for Motion Synthesis and Comprehension

2024-11-26 · Zeyu Ling, Bo Han, Shiyang Li, Hongdeng Shen 외

This paper introduces MotionLLaMA, a unified framework for motion synthesis and comprehension, along with a novel full-body motion tokenizer called the HoMi Tokenizer. MotionLLaMA is developed based on three core princip…

Language ModelingLanguage ModellingLarge Language ModelMotion Synthesis+1

SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling

2026-08-25 · Tianlv Huang, Hetian Guo, Ziyi Cai, Song Wang 외 arxiv

Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for reconstruction and do not explicitly allocate capacity according to …

Large Motion Model for Unified Multi-Modal Motion Generation

2024-04-01 · Mingyuan Zhang, Daisheng Jin, Chenyang Gu, Fangzhou Hong 외

Human motion generation, a cornerstone technique in animation and video production, has widespread applications in various tasks like text-to-motion and music-to-dance. Previous works focus on developing specialist model…

Motion Generation

M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation

2024-05-25 · Mingshuang Luo, Ruibing Hou, Zhuo Li, Hong Chang 외

This paper presents M$^3$GPT, an advanced $\textbf{M}$ultimodal, $\textbf{M}$ultitask framework for $\textbf{M}$otion comprehension and generation. M$^3$GPT operates on three fundamental principles. The first focuses on …

Language ModelingLanguage ModellingLarge Language ModelMotion Generation+2

MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities

2025-04-03 · CVPR 2025 1 · Bizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong 외

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where …

Language ModelingLanguage Modelling