paper-with-me

Papers

Evaluating VLMs' Spatial Reasoning Over Robot Motion: A Step Towards Robot Planning with Motion Preferences

2026-03-13 · Wenxi Wu, Jingjing Zhang, Martim Brandão arxiv

Understanding user instructions and object spatial relations in surrounding environments is crucial for intelligent robot systems to assist humans in various tasks. The natural language and spatial reasoning capabilities of Vision-Language Models (VLMs) have the potential to enhance the generalization of robot planners on new tasks, objects, and motion specifications. While foundation models have been applied to task planning, it is still unclear the degree to which they have the capability of spatial reasoning required to enforce user preferences or constraints on motion, such as desired distances from objects, topological properties, or motion style preferences. In this paper, we evaluate the capability of four state-of-the-art VLMs at spatial reasoning over robot motion, using four different querying methods. Our results show that, with the highest-performing querying method, Qwen2.5-VL achieves 71.4% accuracy zero-shot and 75% on a smaller model after fine-tuning, and GPT-4o leads to lower performance. We evaluate two types of motion preferences (object-proximity and path-style), and we also analyze the trade-off between accuracy and computation cost in number of tokens. This work shows some promise in the potential of VLM integration with robot motion planning pipelines.

📄 PDF Abstract BibTeX arXiv:2603.13100

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningMotion Planning

Similar Papers 제목 키워드 기반

SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

2024-06-03 · An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo 외

Vision Language Models (VLMs) have demonstrated remarkable performance in 2D vision and language tasks. However, their ability to reason about spatial arrangements remains limited. In this work, we introduce Spatial Regi…

Language ModellingSpatial Reasoning

Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models

2025-10-15 · Xinmiao Huang, Qisong He, Zhenglin Huang, Boxuan Wang 외 arxiv

Spatial reasoning ability is crucial for Vision Language Models (VLMs) to support real-world applications in diverse domains including robotics, augmented reality, and autonomous navigation. Unfortunately, existing bench…

Spatial Reasoning

SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation

2025-09-10 · Michael J. Munje, Chen Tang, Shuijing Liu, Zichao Hu 외 arxiv

Robot navigation in dynamic, human-centered environments requires socially-compliant decisions grounded in robust scene understanding. Recent Vision-Language Models (VLMs) exhibit promising capabilities such as object re…

Visual Question AnsweringScene UnderstandingObject RecognitionRobot Navigation

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

2025-10-22 · Zhiyuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du 외 arxiv

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet …

Spatial Reasoning

ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models

2026-03-13 · Yanpeng Zhao, Wentao Ding, Hongtao Li, Baoxiong Jia 외 arxiv

A recent trend in vision-language models (VLMs) has been to enhance their spatial cognition for embodied domains. Despite progress, existing evaluations have been limited both in paradigm and in coverage, hindering rapid…

Question AnsweringSpatial Reasoning