Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning
Vision-language models excel at 2D visual understanding but remain limited in 3D spatial reasoning. Existing approaches either depend on explicit 3D modalities, which limits scalability, or inject partial, view-conditioned geometric priors and leave the language model to recover global scene structure from sparse cues. We introduce Spa3R, a self-supervised framework that learns a unified, view-invariant spatial representation from unposed multi-view RGB images. Its Predictive Spatial Field Modeling objective compresses context views into a compact latent representation and predicts aligned geometric and semantic feature fields at novel viewpoints, thereby encouraging coherent encoding of scene geometry and layout. We integrate the pre-trained Spa3R Encoder into a vision-language model through a lightweight residual cross-attention adapter, yielding Spa3-VLM and grounding language reasoning in global spatial context. Spa3-VLM achieves an average score of 58.6% on VSI-Bench and delivers leading or competitive performance across three additional spatial reasoning benchmarks. These results demonstrate that predictive spatial representation learning provides an effective visual foundation for 3D reasoning.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual ReasoningSimilar Papers 제목 키워드 기반
A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding
Open-vocabulary 3D visual grounding aims to localize target objects based on free-form language queries, which is crucial for embodied AI applications such as autonomous navigation, robotics, and augmented reality. Learn…
3D visual groundingAutonomous NavigationLarge Language ModelSpatial Reasoning+1Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an int…
STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation
Robotic manipulation requires reasoning about future spatial-temporal interactions and geometric constraints, yet existing Vision-Language-Action (VLA) policies often leave predictive representation weakly coupled with a…
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts
Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchmark aimed at assessing the capabilities …
Decision MakingLogical ReasoningQuestion AnsweringSpatial Reasoning+2Improving Vision-and-Language Reasoning via Spatial Relations Modeling
Visual commonsense reasoning (VCR) is a challenging multi-modal task, which requires high-level cognition and commonsense reasoning ability about the real world. In recent years, large-scale pre-training approaches have …
Position regressionRelationRelation ClassificationVisual Commonsense Reasoning+1