paper-with-me

홈 › Papers

MotionLLM: Understanding Human Behaviors from Human Motions and Videos

2024-05-30 · Ling-Hao Chen, Shunlin Lu, Ailing Zeng, Hao Zhang, Benyou Wang, Ruimao Zhang, Lei Zhang

This study delves into the realm of multi-modality (i.e., video and motion modalities) human behavior understanding by leveraging the powerful capabilities of Large Language Models (LLMs). Diverging from recent LLMs designed for video-only or motion-only understanding, we argue that understanding human behavior necessitates joint modeling from both videos and motion sequences (e.g., SMPL sequences) to capture nuanced body part dynamics and semantics effectively. In light of this, we present MotionLLM, a straightforward yet effective framework for human motion understanding, captioning, and reasoning. Specifically, MotionLLM adopts a unified video-motion training strategy that leverages the complementary advantages of existing coarse video-text data and fine-grained motion-text data to glean rich spatial-temporal insights. Furthermore, we collect a substantial dataset, MoVid, comprising diverse videos, motions, captions, and instructions. Additionally, we propose the MoVid-Bench, with carefully manual annotations, for better evaluation of human behavior understanding on video and motion. Extensive experiments show the superiority of MotionLLM in the caption, spatial-temporal comprehension, and reasoning ability.

📄 PDF Abstract BibTeX arXiv:2405.20340

Code (1)

IDEA-Research/MotionLLM 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs

2024-05-27 · Qi Wu, Yubo Zhao, Yifan Wang, Xinhang Liu 외

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an e…

Language ModelingLanguage ModellingMotion CaptioningMotion Generation

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation

2025-12-11 · Yuan-Ming Li, Qize Yang, Nan Lei, Shenghao Fu 외 arxiv

Recent advances in motion-aware large language models have shown remarkable promise for jointly learning motion understanding and generation knowledge. However, these models typically treat understanding and generation s…

Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound

2025-11-15 · Dengming Zhang, Weitao You, Jingxiong Li, Weishen Lin 외 arxiv

Emotion understanding is critical for making Large Language Models (LLMs) more general, reliable, and aligned with humans. Art conveys emotion through the joint design of visual and auditory elements, yet most prior work…

MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities

2025-04-03 · CVPR 2025 1 · Bizhu Wu, Jinheng Xie, Keming Shen, Zhe Kong 외

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where …

Language ModelingLanguage Modelling

HADREB: Human Appraisals and (English) Descriptions of Robot Emotional Behaviors

2022-06-01 · LREC 2022 6 · Josue Torres-Fonsesca, Casey Kennington

Humans sometimes anthropomorphize everyday objects, but especially robots that have human-like qualities and that are often able to interact with and respond to humans in ways that other objects cannot. Humans especially…

AttributeLanguage Modelling