paper-with-me

홈 › Papers

MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities

2025-04-03 · CVPR 2025 1 · Bizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong, Jianfeng Ren, Ruibin Bai, Rong Qu, Linlin Shen

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM

📄 PDF Abstract BibTeX arXiv:2504.02478

Code (1)

cvi-szu/mg-motionllm 공식 구현

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

Focus 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

MotionLLM: Understanding Human Behaviors from Human Motions and Videos

2024-05-30 · Ling-Hao Chen, Shunlin Lu, Ailing Zeng, Hao Zhang 외

This study delves into the realm of multi-modality (i.e., video and motion modalities) human behavior understanding by leveraging the powerful capabilities of Large Language Models (LLMs). Diverging from recent LLMs desi…

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation

2025-12-11 · Yuan-Ming Li, Qize Yang, Nan Lei, Shenghao Fu 외 arxiv

Recent advances in motion-aware large language models have shown remarkable promise for jointly learning motion understanding and generation knowledge. However, these models typically treat understanding and generation s…

Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs

2024-05-27 · Qi Wu, Yubo Zhao, Yifan Wang, Xinhang Liu 외

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an e…

Language ModelingLanguage ModellingMotion CaptioningMotion Generation

MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing

2025-09-28 · Ruibing Hou, Mingshuang Luo, Hongyu Pan, Hong Chang 외 arxiv

This paper proposes MotionVerse, a unified framework that harnesses the capabilities of Large Language Models (LLMs) to comprehend, generate, and edit human motion in both single-person and multi-person scenarios. To eff…

Computational Efficiency

Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound

2025-11-15 · Dengming Zhang, Weitao You, Jingxiong Li, Weishen Lin 외 arxiv

Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual and auditory elements, yet most prior work…