paper-with-me

홈 › Papers

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

2026-07-01 · Xudong Li, Mengdan Zhang, Peixian Chen, Jiaxi Tan, Zihao Huang, Jingyuan Zheng, Yan Zhang, Xiawu Zheng, Xing Sun, Rongrong Ji arxiv

Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations beyond basic object recognition. Existing methods typically build such representations through textual reasoning or 3D reconstruction. However, they often falter during multi-step reasoning, particularly when required to dynamically re-anchor evidence to the specific camera-, object-, or direction-centric reference frames demanded by complex queries. To address this, we propose OmniView-Space, a framework designed to maintain spatial consistency through multimodal egocentric evidence. Our approach consists of three core components: (1) Multi-Perspective Spatial Mapping (MPSM), which re-anchors reconstructed geometry into a query-aligned visual cognitive map and a textual spatial graph; (2) Tool-Guided Egocentric Reasoning, an interleaved policy trained to actively select the ego anchor required by the query and request the corresponding MPSM evidence; and (3) Cognitive-Map Distillation, which uses MPSM-generated trajectories and ego-frame rewards to train the model to reason with self-generated cognitive maps. Experiments on single- and multi-image spatial reasoning benchmarks show that OmniView-Space achieves state-of-the-art performance. Furthermore, the distilled model maintains this performance while reducing reliance on external geometry pipelines.

📄 PDF Abstract BibTeX arXiv:2607.00881

Code (0)

등록된 구현이 없습니다.

Tasks

Object RecognitionSpatial Reasoning3D Reconstruction

Similar Papers 제목 키워드 기반

OmniView: An All-Seeing Diffusion Model for 3D and 4D View Synthesis

2025-12-11 · Xiang Fan, Sharath Girish, Vivek Ramanujan, Chaoyang Wang 외 arxiv

Prior approaches injecting camera control into diffusion models have focused on specific subsets of 4D consistency tasks: novel view synthesis, text-to-video with camera control, image-to-video, amongst others. Therefore…

Novel View SynthesisVideo Generation

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

2025-04-02 · Kun Ouyang, Yuanxin Liu, HaoNing Wu, Yi Liu 외

Video spatial reasoning, which involves inferring the underlying spatial structure from observed video frames, poses a significant challenge for existing Multimodal Large Language Models (MLLMs). This limitation stems pr…

MMESpatial ReasoningVideo MMEVideo Understanding

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

2025-06-11 · Junfei Wu, Jian Guan, Kaituo Feng, Qiang Liu 외

As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, exis…

Multimodal ReasoningSpatial Reasoning

VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning

2025-09-29 · Zhaozhi Wang, Tong Zhang, Mingyue Guo, Yaowei Wang 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved impressive progress in vision-language alignment, yet they remain limited in visual-spatial reasoning. We first identify that this limitation arises from the attenti…

Spatial ReasoningVisual Grounding

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

2026-09-03 · Yijun Yang, Shenghe Zheng, Wenbo Li, Jianhui Liu 외 hf

Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a …

Reinforcement LearningSpatial Reasoning