paper-with-me

홈 › Papers

NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving

2025-04-04 · Kexin Tian, Jingrui Mao, Yunlong Zhang, Jiwan Jiang, Yang Zhou, Zhengzhong Tu

Recent advancements in Vision-Language Models (VLMs) have demonstrated strong potential for autonomous driving tasks. However, their spatial understanding and reasoning-key capabilities for autonomous driving-still exhibit significant limitations. Notably, none of the existing benchmarks systematically evaluate VLMs' spatial reasoning capabilities in driving scenarios. To fill this gap, we propose NuScenes-SpatialQA, the first large-scale ground-truth-based Question-Answer (QA) benchmark specifically designed to evaluate the spatial understanding and reasoning capabilities of VLMs in autonomous driving. Built upon the NuScenes dataset, the benchmark is constructed through an automated 3D scene graph generation pipeline and a QA generation pipeline. The benchmark systematically evaluates VLMs' performance in both spatial understanding and reasoning across multiple dimensions. Using this benchmark, we conduct extensive experiments on diverse VLMs, including both general and spatial-enhanced models, providing the first comprehensive evaluation of their spatial capabilities in autonomous driving. Surprisingly, the experimental results show that the spatial-enhanced VLM outperforms in qualitative QA but does not demonstrate competitiveness in quantitative QA. In general, VLMs still face considerable challenges in spatial understanding and reasoning.

📄 PDF Abstract BibTeX arXiv:2504.03164

Code (0)

등록된 구현이 없습니다.

Tasks

3d scene graph generationAutonomous DrivingGraph GenerationScene Graph GenerationSpatial Reasoning

Similar Papers 제목 키워드 기반

SpatiaLQA: A Benchmark for Evaluating Spatial Logical Reasoning in Vision-Language Models

2026-02-24 · Yuechen Xie, Xiaoyan Zhang, Yicheng Shan, Hao Zhu 외 arxiv

Vision-Language Models (VLMs) have been increasingly applied in real-world scenarios due to their outstanding understanding and reasoning capabilities. Although VLMs have already demonstrated impressive capabilities in c…

Visual Question AnsweringLogical Reasoning

SpatialBot: Precise Spatial Understanding with Vision Language Models

2024-06-19 · Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li 외

Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding, however they are still struggling with spatial understanding which is the foundation of Embodied AI. In this paper, we propose…

Spatial Reasoning

SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning

2026-01-10 · Chenxu Dang, Jie Wang, Guang Li, Zhiwen Hou 외 arxiv

In autonomous driving, Vision Language Models (VLMs) excel at high-level reasoning , whereas semantic occupancy provides fine-grained details. Despite significant progress in individual fields, there is still no method t…

Scene UnderstandingTrajectory PlanningAutonomous Driving

Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving

2025-11-24 · Jianhua Han, Meng Tian, Jiangtong Zhu, Fan He 외 arxiv

Autonomous driving heavily relies on accurate and robust spatial perception. Many failures arise from inaccuracies and instability, especially in long-tail scenarios and complex interactions. However, current vision-lang…

Scene UnderstandingAutonomous DrivingSpatial ReasoningLogical Reasoning

Embodied Scene Understanding for Vision Language Models via MetaVQA

2025-01-15 · CVPR 2025 1 · Weizhen Wang, Chenda Duan, Zhenghao Peng, Yuxin Liu 외

Vision Language Models (VLMs) demonstrate significant potential as embodied AI agents for various mobility applications. However, a standardized, closed-loop benchmark for evaluating their spatial reasoning and sequentia…

Decision MakingQuestion AnsweringScene UnderstandingSequential Decision Making+3