paper-with-me

Papers

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning

2025-04-28 · Wufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang, Celso de Melo, Jianwen Xie, Alan Yuille

Despite recent advances on multi-modal models, 3D spatial reasoning remains a challenging task for state-of-the-art open-source and proprietary models. Recent studies explore data-driven approaches and achieve enhanced spatial reasoning performance by fine-tuning models on 3D-related visual question-answering data. However, these methods typically perform spatial reasoning in an implicit manner and often fail on questions that are trivial to humans, even with long chain-of-thought reasoning. In this work, we introduce SpatialReasoner, a novel large vision-language model (LVLM) that addresses 3D spatial reasoning with explicit 3D representations shared between multiple stages--3D perception, computation, and reasoning. Explicit 3D representations provide a coherent interface that supports advanced 3D spatial reasoning and improves the generalization ability to novel question types. Furthermore, by analyzing the explicit 3D representations in multi-step reasoning traces of SpatialReasoner, we study the factual errors and identify key shortcomings of current LVLMs. Results show that our SpatialReasoner achieves improved performance on a variety of spatial reasoning benchmarks, outperforming Gemini 2.0 by 9.2% on 3DSRBench, and generalizes better when evaluating on novel 3D spatial reasoning questions. Our study bridges the 3D parsing capabilities of prior visual foundation models with the powerful reasoning abilities of large language models, opening new directions for 3D spatial reasoning.

📄 PDF Abstract BibTeX arXiv:2504.20024

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringSpatial ReasoningVisual Question Answering

Similar Papers 제목 키워드 기반

A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding

2025-07-09 · Zhenyang Liu, Sixiao Zheng, Siyu Chen, Cairong Zhao 외

Open-vocabulary 3D visual grounding aims to localize target objects based on free-form language queries, which is crucial for embodied AI applications such as autonomous navigation, robotics, and augmented reality. Learn…

3D visual groundingAutonomous NavigationLarge Language ModelSpatial Reasoning+1

SpatialReasoner: Active Perception for Large-Scale 3D Scene Understanding

2025-12-02 · Hongpei Zheng, Shijie Li, Yanran Li, Hujun Yin arxiv

Spatial reasoning in large-scale 3D environments remains challenging for current vision-language models, which are typically constrained to room-scale scenarios. We introduce H$^2$U3D (Holistic House Understanding in 3D)…

Visual Question AnsweringReinforcement LearningScene UnderstandingSpatial Reasoning

Spatial Reasoners for Continuous Variables in Any Domain

2025-07-14 · Bart Pogodzinski, Christopher Wewer, Bernt Schiele, Jan Eric Lenssen arxiv

We present Spatial Reasoners, a software framework to perform spatial reasoning over continuous variables with generative denoising models. Denoising generative models have become the de-facto standard for image generati…

Spatial ReasoningImage Generation

SEM: Enhancing Spatial Understanding for Robust Robot Manipulation

2025-05-22 · Xuewu Lin, Tianwei Lin, Lichao Huang, Hongyu Xie 외

A key challenge in robot manipulation lies in developing policy models with strong spatial understanding, the ability to reason about 3D geometry, object relations, and robot embodiment. Existing methods often fall short…

3D geometryRobot ManipulationSpatial Reasoning

Generalizable Operating Room Expert with Multimodal Enhancement

2025-08-11 · Peiqi He, Zhenhao Zhang, Yixiang Zhang, Xiongjun Zhao 외 arxiv

Precise spatial modeling in the operating room (OR) is essential for intraoperative awareness, hazard avoidance, and surgical decision-making. Although existing approaches exploit multimodal data to learn spatial relatio…

Spatial Reasoning