paper-with-me

Papers

Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs

2026-01-14 · Rui Zhu, Xin Shen, Shuchen Wu, Chenxi Miao, Xin Yu, Yang Li, Weikang Li, Deguo Xia, Jizhou Huang arxiv

Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmarks primarily focus on single-step perception-to-judgment tasks, leaving scenarios requiring complex visual-spatial logical chains significantly underexplored. To bridge this gap, we introduce Video-MSR, the first benchmark specifically designed to evaluate Multi-hop Spatial Reasoning (MSR) in dynamic video scenarios. Video-MSR systematically probes MSR capabilities through four distinct tasks: Constrained Localization, Chain-based Reference Retrieval, Route Planning, and Counterfactual Physical Deduction. Our benchmark comprises 3,052 high-quality video instances with 4,993 question-answer pairs, constructed via a scalable, visually-grounded pipeline combining advanced model generation with rigorous human verification. Through a comprehensive evaluation of 20 state-of-the-art MLLMs, we uncover significant limitations, revealing that while models demonstrate proficiency in surface-level perception, they exhibit distinct performance drops in MSR tasks, frequently suffering from spatial disorientation and hallucination during multi-step deductions. To mitigate these shortcomings and empower models with stronger MSR capabilities, we further curate MSR-9K, a specialized instruction-tuning dataset, and fine-tune Qwen-VL, achieving a +7.82% absolute improvement on Video-MSR. Our results underscore the efficacy of multi-hop spatial instruction data and establish Video-MSR as a vital foundation for future research. The code and data will be available at https://github.com/ruiz-nju/Video-MSR.

📄 PDF Abstract BibTeX arXiv:2601.09430

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

2025-09-30 · Shangding Gu, Xiaohan Wang, Donghao Ying, Haoyu Zhao 외 arxiv

Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench, a large-scale benchmark that combines …

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs

2026-05-04 · Le Zhang, Jihan Yang, Soundarya Krishnan, Jimit Majmudar 외 arxiv

Human-level agentic intelligence extends beyond low-level geometric perception, evolving from recognizing where things are to understanding what they are for. While existing benchmarks effectively evaluate the geometric …

Relational ReasoningSpatial Reasoning

HourVideo: 1-Hour Video-Language Understanding

2024-11-07 · Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota 외

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, tempora…

BenchmarkingcounterfactualMultiple-choiceRetrieval+1

Mind over Space: Can Multimodal Large Language Models Mentally Navigate?

2026-03-23 · Qihui Zhu, Shouwei Ruan, Xiao Yang, Hao Jiang 외 arxiv

Despite the widespread adoption of MLLMs in embodied agents, their capabilities remain largely confined to reactive planning from immediate observations, consistently failing in spatial reasoning across extensive spatiot…

Spatial Reasoning

TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models

2025-11-17 · Harold Haodong Chen, Disen Lan, Wen-Jie Shu, Qingyang Liu 외 arxiv

The rapid evolution of video generative models has shifted their focus from producing visually plausible outputs to tackling tasks requiring physical plausibility and logical consistency. However, despite recent breakthr…

Logical ReasoningVideo Generation