paper-with-me

Papers

Fine-grained Human Motion Understanding with Language Models

2026-06-18 · Thomas Markhorst, Zhi-Yi Lin, Jouh Yeong Chew, Jan van Gemert, Xucong Zhang arxiv

In this work, we propose \methodname, an LLM-based model for fine-grained human motion understanding that represents motion as a sequence of skeletal poses with explicit timestamps for each pose. Each pose encodes body joint positions and is temporally grounded with timestamp tokens, allowing the model to reason about motion order, duration, and rhythm. To study what supervision is needed for motion-language reasoning, we construct a diverse training mixture spanning pose captioning, pose question answering, motion captioning, and motion question answering. Our ablations show that the primary gains come from the diversity of pose- and motion-level supervision, while staged training provides a smaller additional benefit. Different from previous works that rely on ground-truth 3D motion capture, our approach supports both 2D and 3D skeletal motion representations through a unified pose encoder, and can optionally incorporate video to provide contextual information. Extensive experiments on BABEL-QA, HuMMan-QA, CompMo, NTU-RGB+D, and QEVD-Coach demonstrate that our method achieves state-of-the-art performance across multiple benchmarks, highlighting the effectiveness of explicit temporal encoding and diverse pose- and motion-level supervision for fine-grained human motion understanding. Notably, even when using only 2D skeletal input, our approach surpasses previous 3D-based methods.

📄 PDF Abstract BibTeX arXiv:2606.20888

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringMotion Captioning

Similar Papers 제목 키워드 기반

Fluent but Unfeeling: The Emotional Blind Spots of Language Models

2025-09-11 · Bangzhao Shu, Isha Joshi, Melissa Karnaze, Anh C. Pham 외 arxiv

The versatility of Large Language Models (LLMs) in natural language understanding has made them increasingly popular in mental health research. While many studies explore LLMs' capabilities in emotion recognition, a crit…

Natural Language UnderstandingEmotion Recognition

MotionMERGE: A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation

2026-05-18 · Bizhu Wu, Jinheng Xie, Wenting Chen, Zhe Kong 외 arxiv

Recent motion-language models unify tasks like comprehension and generation but operate at a coarse granularity, lacking fine-grained understanding and nuanced control over body parts needed for animation or interaction.…

Zero-shot Generalization

SportsCap: Monocular 3D Human Motion Capture and Fine-grained Understanding in Challenging Sports Videos

2021-04-23 · Xin Chen, Anqi Pang, Wei Yang, Yuexin Ma 외

Markerless motion capture and understanding of professional non-daily human movements is an important yet unsolved task, which suffers from complex motion patterns and severe self-occlusion, especially for the monocular …

Action AssessmentAttributeMarkerless Motion Capture

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

2025-01-06 · CVPR 2025 1 · Wenyi Hong, Yean Cheng, Zhuoyi Yang, Weihan Wang 외

In recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability - fine-grained motion comprehension - remains under-explored in current benchmarks. …

BenchmarkingFeature CompressionVideo Understanding

MoChat: Joints-Grouped Spatio-Temporal Grounding LLM for Multi-Turn Motion Comprehension and Description

2024-10-15 · Jiawei Mo, Yixuan Chen, Rifen Lin, Yongkang Ni 외

Despite continuous advancements in deep learning for understanding human motion, existing models often struggle to accurately identify action timing and specific body parts, typically supporting only single-round interac…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model