paper-with-me

Papers

Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning

2026-02-24 · Haoyi Jiang, Liu Liu, Xinjie Wang, Yonghao He, Wei Sui, Zhizhong Su, Wenyu Liu, Xinggang Wang arxiv

Vision-language models excel at 2D visual understanding but remain limited in 3D spatial reasoning. Existing approaches either depend on explicit 3D modalities, which limits scalability, or inject partial, view-conditioned geometric priors and leave the language model to recover global scene structure from sparse cues. We introduce Spa3R, a self-supervised framework that learns a unified, view-invariant spatial representation from unposed multi-view RGB images. Its Predictive Spatial Field Modeling objective compresses context views into a compact latent representation and predicts aligned geometric and semantic feature fields at novel viewpoints, thereby encouraging coherent encoding of scene geometry and layout. We integrate the pre-trained Spa3R Encoder into a vision-language model through a lightweight residual cross-attention adapter, yielding Spa3-VLM and grounding language reasoning in global spatial context. Spa3-VLM achieves an average score of 58.6% on VSI-Bench and delivers leading or competitive performance across three additional spatial reasoning benchmarks. These results demonstrate that predictive spatial representation learning provides an effective visual foundation for 3D reasoning.

📄 PDF Abstract BibTeX arXiv:2602.21186

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding

2025-07-09 · Zhenyang Liu, Sixiao Zheng, Siyu Chen, Cairong Zhao 외

Open-vocabulary 3D visual grounding aims to localize target objects based on free-form language queries, which is crucial for embodied AI applications such as autonomous navigation, robotics, and augmented reality. Learn…

3D visual groundingAutonomous NavigationLarge Language ModelSpatial Reasoning+1

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

2026-08-16 · Xiaoyu Zhu, Xinke Deng, Suresh Taddewadikar, Arnab Kumar Mondal 외 hf

Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an int…

STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation

2026-04-29 · Yuxuan Tian, Yurun Jin, Bin Yu, Yukun Shi 외 arxiv

Robotic manipulation requires reasoning about future spatial-temporal interactions and geometric constraints, yet existing Vision-Language-Action (VLA) policies often leave predictive representation weakly coupled with a…

FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts

2024-06-27 · Shubhankar Singh, Purvi Chaurasia, Yerram Varun, Pranshu Pandya 외

Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. We introduce FlowVQA, a novel benchmark aimed at assessing the capabilities …

Decision MakingLogical ReasoningQuestion AnsweringSpatial Reasoning+2

Improving Vision-and-Language Reasoning via Spatial Relations Modeling

2023-11-09 · Cheng Yang, Rui Xu, Ye Guo, Peixiang Huang 외

Visual commonsense reasoning (VCR) is a challenging multi-modal task, which requires high-level cognition and commonsense reasoning ability about the real world. In recent years, large-scale pre-training approaches have …

Position regressionRelationRelation ClassificationVisual Commonsense Reasoning+1