paper-with-me

Papers

An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models

2024-11-09 · Fatemeh Shiri, Xiao-Yu Guo, Mona Golestan Far, Xin Yu, Gholamreza Haffari, Yuan-Fang Li

Large Multimodal Models (LMMs) have achieved strong performance across a range of vision and language tasks. However, their spatial reasoning capabilities are under-investigated. In this paper, we construct a novel VQA dataset, Spatial-MM, to comprehensively study LMMs' spatial understanding and reasoning capabilities. Our analyses on object-relationship and multi-hop reasoning reveal several important findings. Firstly, bounding boxes and scene graphs, even synthetic ones, can significantly enhance LMMs' spatial reasoning. Secondly, LMMs struggle more with questions posed from the human perspective than the camera perspective about the image. Thirdly, chain of thought (CoT) prompting does not improve model performance on complex multi-hop questions involving spatial relations. % Moreover, spatial reasoning steps are much less accurate than non-spatial ones across MLLMs. Lastly, our perturbation analysis on GQA-spatial reveals that LMMs are much stronger at basic object detection than complex spatial reasoning. We believe our benchmark dataset and in-depth analyses can spark further research on LMMs spatial reasoning. Spatial-MM benchmark is available at: https://github.com/FatemehShiri/Spatial-MM

📄 PDF Abstract BibTeX arXiv:2411.06048

Code (1)

fatemehshiri/spatial-mm 공식 구현

Tasks

object-detectionObject DetectionSpatial ReasoningVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses

2025-08-07 · Bin Han, Robert Wolfe, Anat Caspi, Bill Howe arxiv

We explore the application of large language models (LLMs) to empower domain experts in integrating large, heterogeneous, and noisy urban spatial datasets. Traditional rule-based integration methods are unable to cover a…

Spatial Reasoning

Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models

2025-03-21 · Jianing Qi, Jiawei Liu, Hao Tang, Zhigang Zhu

Vision-Language Models (VLMs) excel at identifying and describing objects but struggle with spatial reasoning such as accurately understanding the relative positions of objects. Inspired by the dual-pathway (ventral-dors…

DiagnosticObject RecognitionSpatial Reasoning

A Call for New Recipes to Enhance Spatial Reasoning in MLLMs

2025-04-21 · Huanyu Zhang, Chengzu Li, Wenshan Wu, Shaoguang Mao 외

Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in general vision-language tasks. However, recent studies have exposed critical limitations in their spatial reasoning capabilities. This …

Spatial Reasoning

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks

2025-09-02 · Nils Hoehing, Mayug Maniparambil, Ellen Rushe, Noel E. O'Connor 외 arxiv

We propose RocketScience, an open-source contrastive VLM benchmark that tests for spatial relation understanding. It is comprised of entirely new real-world image-text pairs covering mostly relative spatial understanding…

Object LocalizationSpatial Reasoning

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis

2025-08-27 · Chengzu Li, Wenshan Wu, Huanyu Zhang, Qingtao Li 외 arxiv

For human cognitive process, spatial reasoning and perception are closely entangled, yet the nature of this interplay remains underexplored in the evaluation of multimodal large language models (MLLMs). While recent MLLM…

Spatial Reasoning