paper-with-me

홈 › Papers

3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

2024-12-10 · Wufei Ma, Haoyu Chen, Guofeng Zhang, Yu-Cheng Chou, Celso M de Melo, Alan Yuille

3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a comprehensive understanding of the 3D scene, enabling their applicability to a broader range of areas, such as autonomous navigation, robotics, and AR/VR. While large multi-modal models (LMMs) have achieved remarkable progress in a wide range of image and video understanding tasks, their capabilities to perform 3D spatial reasoning on diverse natural images are less studied. In this work we present the first comprehensive 3D spatial reasoning benchmark, 3DSRBench, with 2,772 manually annotated visual question-answer pairs across 12 question types. We conduct robust and thorough evaluation of 3D spatial reasoning capabilities by balancing the data distribution and adopting a novel FlipEval strategy. To further study the robustness of 3D spatial reasoning w.r.t. camera 3D viewpoints, our 3DSRBench includes two subsets with 3D spatial reasoning questions on paired images with common and uncommon viewpoints. We benchmark a wide range of open-sourced and proprietary LMMs, uncovering their limitations in various aspects of 3D awareness, such as height, orientation, location, and multi-object reasoning, as well as their degraded performance on images with uncommon camera viewpoints. Our 3DSRBench provide valuable findings and insights about the future development of LMMs with strong 3D reasoning capabilities. Our project page and dataset is available https://3dsrbench.github.io.

📄 PDF Abstract BibTeX arXiv:2412.07825

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous NavigationSpatial ReasoningVideo Understanding

Similar Papers 제목 키워드 기반

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning

2025-04-28 · Wufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang 외

Despite recent advances on multi-modal models, 3D spatial reasoning remains a challenging task for state-of-the-art open-source and proprietary models. Recent studies explore data-driven approaches and achieve enhanced s…

Question AnsweringSpatial ReasoningVisual Question Answering

Cognitively-Inspired Tokens Overcome Egocentric Bias in Multimodal Models

2026-01-23 · Bridget Leonard, Scott O. Murray arxiv

Multimodal language models (MLMs) perform well on semantic vision-language tasks but fail at spatial reasoning that requires adopting another agent's visual perspective. These errors reflect a persistent egocentric bias …

Spatial Reasoning

SpatiO: Adaptive Test-Time Orchestration of Vision-Language Agents for Spatial Reasoning

2026-04-23 · Chan Yeong Hwang, Miso Choi, Sunghyun On, Jinkyu Kim 외 arxiv

Understanding visual scenes requires not only recognizing objects but also reasoning about their spatial relationships. Unlike general vision-language tasks, spatial reasoning requires integrating multiple inductive bias…

Spatial Reasoning

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning

2026-04-08 · Jiahua Chen, Qihong Tang, Weinong Wang, Qi Fan arxiv

Although Multimodal Large Language Models have achieved remarkable progress, they still struggle with complex 3D spatial reasoning due to the reliance on 2D visual priors. Existing approaches typically mitigate this limi…

Keyword ExtractionSpatial Reasoning3D Reconstruction

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

2025-06-03 · Mengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang 외

Spatial reasoning is a key aspect of cognitive psychology and remains a major bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic s…

Object CountingSpatial Reasoning