paper-with-me

홈 › Papers

MotionLLaMA: A Unified Framework for Motion Synthesis and Comprehension

2024-11-26 · Zeyu Ling, Bo Han, Shiyang Li, Hongdeng Shen, Jikang Cheng, Changqing Zou

This paper introduces MotionLLaMA, a unified framework for motion synthesis and comprehension, along with a novel full-body motion tokenizer called the HoMi Tokenizer. MotionLLaMA is developed based on three core principles. First, it establishes a powerful unified representation space through the HoMi Tokenizer. Using a single codebook, the HoMi Tokenizer in MotionLLaMA achieves reconstruction accuracy comparable to residual vector quantization tokenizers utilizing six codebooks, outperforming all existing single-codebook tokenizers. Second, MotionLLaMA integrates a large language model to tackle various motion-related tasks. This integration bridges various modalities, facilitating both comprehensive and intricate motion synthesis and comprehension. Third, MotionLLaMA introduces the MotionHub dataset, currently the most extensive multimodal, multitask motion dataset, which enables fine-tuning of large language models. Extensive experimental results demonstrate that MotionLLaMA not only covers the widest range of motion-related tasks but also achieves state-of-the-art (SOTA) performance in motion completion, interaction dual-person text-to-motion, and all comprehension tasks while reaching performance comparable to SOTA in the remaining tasks. The code and MotionHub dataset are publicly available.

📄 PDF Abstract BibTeX arXiv:2411.17335

Code (1)

ZeyuLing/MotionLLaMA 공식 구현 pytorch

Tasks

Language ModelingLanguage ModellingLarge Language ModelMotion SynthesisQuantization

Similar Papers 제목 키워드 기반

Human Motion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy

2025-11-11 · Gong Jingyu, Tong Kunkun, Chen Zhuoran, Yuan Chuanhan 외 arxiv

Human motion synthesis in 3D scenes relies heavily on scene comprehension, while current methods focus mainly on scene structure but ignore the semantic understanding. In this paper, we propose a human motion synthesis f…

Dimensionality ReductionMotion Synthesis

MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities

2025-04-03 · CVPR 2025 1 · Bizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong 외

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where …

Language ModelingLanguage Modelling

MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing

2025-09-28 · Ruibing Hou, Mingshuang Luo, Hongyu Pan, Hong Chang 외 arxiv

This paper proposes MotionVerse, a unified framework that harnesses the capabilities of Large Language Models (LLMs) to comprehend, generate, and edit human motion in both single-person and multi-person scenarios. To eff…

Computational Efficiency

MUSE: Manipulating Unified Framework for Synthesizing Emotions in Images via Test-Time Optimization

2025-11-26 · Yingjie Xia, Xi Wang, Jinglei Shi, Vicky Kalogeiton 외 arxiv

Images evoke emotions that profoundly influence perception, often prioritized over content. Current Image Emotional Synthesis (IES) approaches artificially separate generation and editing tasks, creating inefficiencies a…

Semantic Similarity

FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis

2024-05-24 · Ke Fan, Junshu Tang, Weijian Cao, Ran Yi 외

Text-to-motion synthesis is a crucial task in computer vision. Existing methods are limited in their universality, as they are tailored for single-person or two-person scenarios and can not be applied to generate motions…

Motion GenerationMotion Synthesis