paper-with-me

Papers

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

2025-01-06 · CVPR 2025 1 · Wenyi Hong, Yean Cheng, Zhuoyi Yang, Weihan Wang, Lefan Wang, Xiaotao Gu, Shiyu Huang, Yuxiao Dong, Jie Tang

In recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability - fine-grained motion comprehension - remains under-explored in current benchmarks. To address this gap, we propose MotionBench, a comprehensive evaluation benchmark designed to assess the fine-grained motion comprehension of video understanding models. MotionBench evaluates models' motion-level perception through six primary categories of motion-oriented question types and includes data collected from diverse sources, ensuring a broad representation of real-world video content. Experimental results reveal that existing VLMs perform poorly in understanding fine-grained motions. To enhance VLM's ability to perceive fine-grained motion within a limited sequence length of LLM, we conduct extensive experiments reviewing VLM architectures optimized for video feature compression and propose a novel and efficient Through-Encoder (TE) Fusion method. Experiments show that higher frame rate inputs and TE Fusion yield improvements in motion understanding, yet there is still substantial room for enhancement. Our benchmark aims to guide and motivate the development of more capable video understanding models, emphasizing the importance of fine-grained motion comprehension. Project page: https://motion-bench.github.io .

📄 PDF Abstract BibTeX arXiv:2501.02955

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingFeature CompressionVideo Understanding

Similar Papers 제목 키워드 기반

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

2025-03-19 · Chongjun Tu, Lin Zhang, Pengtao Chen, Peng Ye 외

Multimodal Large Language Models (MLLMs) have shown remarkable capabilities in video content understanding but still struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding abi…

BenchmarkingMultiple-choiceVideo Understanding

Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning

2025-06-05 · Yue Ma, Yulong Liu, Qiyuan Zhu, Ayden Yang 외

Recently, breakthroughs in the video diffusion transformer have shown remarkable capabilities in diverse motion generations. As for the motion-transfer task, current methods mainly use two-stage Low-Rank Adaptations (LoR…

Efficient Motion-Aware Video MLLM

2025-01-01 · CVPR 2025 1 · Zijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo 외

Most current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Mot…

Question AnsweringVideo Question AnsweringVideo Understanding

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning

2026-06-03 · Haozhe Chi, Yang Jin, Yadong Mu arxiv

Despite significant progress in agentic long video understanding, existing methods still lack detailed motion comprehension coupled with an efficient memory architecture. In this paper, we propose GOPAgen, a novel approa…

Video Question Answering

MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction

2026-06-17 · Jianing Zhang, Chenhao Zheng, Yajun Yang, Max Argus 외 arxiv

Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan actions, reason about physical interactions, and synthesize realistic futures. We argue that 3D points i…

Robot ManipulationMotion Forecasting