paper-with-me

홈 › Papers

HumanMoveVQA: Can Video MLLMs reason about human movement in videos?

2026-06-26 · Pulkit Gera, Faegheh Sardari, Asmar Nadeem, Valentina Bono, Padraig Boulton, Adrian Hilton, Armin Mustafa arxiv

Despite the rapid advance of Multimodal Large Language Models (MLLMs) in high-level video understanding, a fundamental bottleneck remains: these models collapse complex human motion into coarse semantic labels. Existing benchmarks mostly focus on scene-centric events or local joint articulations, failing to probe global human motion in space over time (trajectory and orientation changes). We introduce HumanMoveVQA, the first comprehensive benchmark designed to evaluate global trajectory and orientation reasoning from an exocentric perspective. Our benchmark utilizes a first-frame anchored world coordinate system, preserving translation and rotation relative to a fixed starting point. We propose a scalable, multi-stage pipeline that lifts 2D video observations into world-consistent 3D motion tracks to generate over 10K structured question-answer pairs across seven reasoning categories, including motion aggregation, sequential ordering, and trajectory-level inference. Our extensive evaluation reveals a critical capability gap in state-of-the-art proprietary models on deep human motion understanding. However, we demonstrate that this is a learnable problem; by fine-tuning an open-source baseline with our targeted, world-consistent supervision, we achieve a significant improvement. HumanMoveVQA establishes a rigorous geometric foundation for developing next-generation, movement-aware video understanding models.

📄 PDF Abstract BibTeX arXiv:2606.27999

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos

2024-06-12 · Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu 외

Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models" -- interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal m…

counterfactualFuture predictionVideo Understanding

VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning

2025-05-18 · Qi Wang, Yanrui Yu, Ye Yuan, Rui Mao 외

Reinforcement fine-tuning (RFT) has shown great promise in achieving humanlevel reasoning capabilities of Large Language Models (LLMs), and has recently been extended to MLLMs. Nevertheless, reasoning about videos, which…

Reinforcement Learning (RL)

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

2026-05-18 · Yuqi Tang, Yang Shi, Zhuoran Zhang, Qixun Wang 외 arxiv

Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While…

Video ClassificationArtifact Detection

SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models

2024-10-11 · Haotian Xia, Zhengbang Yang, Junbo Zou, Rhys Tracy 외

Multimodal Large Language Models (MLLMs) are advancing the ability to reason about complex sports scenarios by integrating textual and visual information. To comprehensively evaluate their capabilities, we introduce SPOR…

Few-Shot LearningMultiple-choiceQuestion AnsweringSports Understanding

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

2026-06-01 · Xiaolin Liu, Yilun Zhu, Xiangyu Zhao, Xuehui Wang 외 arxiv

Video multimodal large language models (MLLMs) have made rapid progress on general and long-form video understanding, yet their ability to preserve brief answer-critical visual evidence remains underexplored. Many practi…