paper-with-me

홈 › Papers

CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments

2026-01-20 · Haotian Xu, Yue Hu, Zhengqiu Zhu, Chen Gao, Ziyou Wang, Junreng Rao, Wenhao Lu, Weishi Li, Quanjun Yin, Yong Li arxiv

Cross-view spatial reasoning is essential for embodied AI, underpinning spatial understanding, mental simulation and planning in complex environments. Existing benchmarks primarily emphasize indoor or street settings, overlooking the unique challenges of open-ended urban spaces characterized by rich semantics, complex geometries, and view variations. To address this, we introduce CityCube, a systematic benchmark designed to probe cross-view reasoning capabilities of current VLMs in urban settings. CityCube integrates four viewpoint dynamics to mimic camera movements and spans a wide spectrum of perspectives from multiple platforms, e.g., vehicles, drones and satellites. For a comprehensive assessment, it features 5,022 meticulously annotated multi-view QA pairs categorized into five cognitive dimensions and three spatial relation expressions. A comprehensive evaluation of 33 VLMs reveals a significant performance disparity with humans: even large-scale models struggle to exceed 54.1% accuracy, remaining 34.2% below human performance. By contrast, small-scale fine-tuned VLMs achieve over 60.0% accuracy, highlighting the necessity of our benchmark. Further analyses indicate the task correlations and fundamental cognitive disparity between VLMs and human-like reasoning.

📄 PDF Abstract BibTeX arXiv:2601.14339

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

2025-10-22 · Zhiyuan Feng, Zhaolu Kang, Qijie Wang, Zhiying Du 외 arxiv

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet …

Spatial Reasoning

G$^2$TAM: Geometry Grounded Track Anything Model

2026-07-04 · Chenming Zhu, Peizhou Cao, Jingli Lin, Wenbo Hu 외 arxiv

Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explic…

Video Object SegmentationVideo SegmentationSpatial Reasoning3D Reconstruction

Learning Multi-View Spatial Reasoning from Cross-View Relations

2026-03-30 · Suchae Jeong, Jaehwi Song, Haeone Lee, Hanna Kim 외 arxiv

Vision-language models (VLMs) have achieved impressive results on single-view vision tasks, but lack the multi-view spatial reasoning capabilities essential for embodied AI systems to understand 3D environments and manip…

Spatial Reasoning

SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion

2026-06-26 · Haoyu Zhang, Meng Liu, Qianlong Xiang, Kun Wang 외 arxiv

Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation. However, existing UAV benchmarks often emphasize image-level recognition, single-view understand…

MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos

2026-07-01 · Leyuan Yu, Xiao Tang, Minghao Liu, Xinyuan Li 외 arxiv

Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically vary the observer while keeping the scene…

Spatial Reasoning