paper-with-me

Papers

Towards Fine-Grained Human Motion Video Captioning

2025-10-24 · Guorui Song, Guocun Wang, Zhe Huang, Jing Lin, Xuefei Zhe, Jian Li, Haoqian Wang arxiv

Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motion details, resulting in vague or semantically inconsistent captions. In this work, we introduce the Motion-Augmented Caption Model (M-ACM), a novel generative framework that enhances caption quality by incorporating motion-aware decoding. At its core, M-ACM leverages motion representations derived from human mesh recovery to explicitly highlight human body dynamics, thereby reducing hallucinations and improving both semantic fidelity and spatial alignment in the generated captions. To support research in this area, we present the Human Motion Insight (HMI) Dataset, comprising 115K video-description pairs focused on human movement, along with HMI-Bench, a dedicated benchmark for evaluating motion-focused video captioning. Experimental results demonstrate that M-ACM significantly outperforms previous methods in accurately describing complex human motions and subtle temporal variations, setting a new standard for motion-centric video captioning.

📄 PDF Abstract BibTeX arXiv:2510.24767

Code (0)

등록된 구현이 없습니다.

Tasks

Human Mesh RecoveryVideo Captioning

Similar Papers 제목 키워드 기반

MotionAtlas: Detailed Region Captioning for Motion-Centric Videos

2026-06-28 · Weisong Liu, Haochen Wang, Kuan Gao, Yuhao Wang 외 arxiv

We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2) a scalable, high-quality pipeline to construct training samples, and (3) a fami…

Motion Captioning

FingerCap: Fine-grained Finger-level Hand Motion Captioning

2025-11-21 · Xin Shen, Rui Zhu, Lei Shen, Xinyu Wang 외 arxiv

Understanding fine-grained human hand motion is fundamental to visual perception, embodied intelligence, and multimodal communication. In this work, we propose Fine-grained Finger-level Hand Motion Captioning (FingerCap)…

Motion Captioning

Towards Accurate Emotion-Attributed Video Captioning via Fine-grained Emotion-Cause Pair Extraction

2026-06-07 · Weidong Chen, Cheng Ye, Zhendong Mao, Liping Wang 외 arxiv

Emotional Video Captioning (EVC) is a challenging task that aims to generate factually accurate and emotionally rich descriptions for videos. Existing EVC methods leverage holistic visual features to mine global emotiona…

Emotion-Cause Pair ExtractionVideo Captioning

KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding

2026-02-19 · Boda Lin, Yongjie Zhu, Xiaocheng Gong, Wenyu Qin 외 arxiv

Despite recent advancements, video captioning models still face significant limitations in accurately describing fine-grained motion details and suffer from severe hallucination issues. These challenges become particular…

Video Captioning

Fine-grained Human Motion Understanding with Language Models

2026-06-18 · Thomas Markhorst, Zhi-Yi Lin, Jouh Yeong Chew, Jan van Gemert 외 arxiv

In this work, we propose \methodname, an LLM-based model for fine-grained human motion understanding that represents motion as a sequence of skeletal poses with explicit timestamps for each pose. Each pose encodes body j…

Question AnsweringMotion Captioning