paper-with-me

홈 › Papers

Human Motion Instruction Tuning

2024-11-25 · CVPR 2025 1 · Lei LI, Sen Jia, Jianhao Wang, Zhongyu Jiang, Feng Zhou, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang

This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model's ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Our code and models are available on the project website: https://github.com/ILGLJ/LLaMo.

📄 PDF Abstract BibTeX arXiv:2411.16805

Code (0)

등록된 구현이 없습니다.

Tasks

Sports Analytics

Similar Papers 제목 키워드 기반

HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA

2025-11-03 · Lei Hu, Yongjing Ye, Shihong Xia arxiv

The expansion of instruction-tuning data has enabled foundation language models to exhibit improved instruction adherence and superior performance across diverse downstream tasks. Semantically-rich 3D human motion is bei…

RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control

2025-06-15 · Junpeng Yue, Zepeng Wang, Yuxuan Wang, Weishuai Zeng 외

This paper focuses on a critical challenge in robotics: translating text-driven human motions into executable actions for humanoid robots, enabling efficient and cost-effective learning of new behaviors. While existing t…

Humanoid ControlMotion GenerationSemantic correspondence

EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning

2024-04-25 · CVPR 2024 1 · HongXia Xie, Chu-Jun Peng, Yu-Wen Tseng, Hung-Jen Chen 외

Visual Instruction Tuning represents a novel learning paradigm involving the fine-tuning of pre-trained language models using task-specific instructions. This paradigm shows promising zero-shot results in various natural…

Emotion ClassificationEmotion Recognition

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

2025-07-21 · Hao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng 외 arxiv

We introduce Being-H0, a dexterous Vision-Language-Action model (VLA) trained on large-scale human videos. Existing VLAs struggle with complex manipulation tasks requiring high dexterity and generalize poorly to novel sc…

Instruction Following

In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning

2025-08-16 · Hui Ma, Bo Zhang, Jinpeng Hu, Zenglin Shi arxiv

Emotion recognition in conversation (ERC) aims to identify the emotion of each utterance in a conversation, playing a vital role in empathetic artificial intelligence. With the growing of large language models (LLMs), in…

Emotion Recognition in Conversation