paper-with-me

홈 › Papers

Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World

2026-09-19 · Kaixiang Yao, Xu Wang, Miao Pan, Hu Xiyue, Weishi Wang, Daniel Dahlmeier, Jintao Chen, Yongliang Shen, Xuhong Zhang, Wenqi Zhang hf

Spatial reasoning is essential for vision-language models (VLMs) to understand and act in the physical world. Reasoning in dynamic environments requires VLMs to perceive local state transitions caused by object motion and viewpoint changes and integrate them over long trajectories to maintain an updated spatial state, yet existing VLMs remain limited in both capabilities. Current spatial training primarily focuses on static questions about object attributes and spatial relations, providing limited direct supervision for state transitions; in contrast, interaction trajectories naturally connect a preceding observation, an action, and a subsequent observation, offering direct supervision for local state transitions, while complete trajectories reveal dependencies among consecutive transitions. We therefore introduce Spatial-Interactor, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories. Accordingly, we construct the Learning from Spatial Interaction dataset (LSI-108K) from simulated and real interaction trajectories, with tasks aligned with the objective of each level. Our two-stage training strategy applies Supervised Fine-Tuning (SFT) to L1 and L2 for local transition modeling, and On-Policy Distillation (OPD) then uses privileged self-distillation: a teacher branch given segment-level transition descriptions supervises the student's on-policy CoT, helping the student learn to integrate consecutive transitions over L3 long trajectories. Experiments across multiple VLMs and spatial benchmarks show consistent gains in local transition modeling and long-horizon integration.

📄 PDF Abstract BibTeX arXiv:2609.23038

Code (1)

ZJU-OmniAI/Spatial-Interactor ★ 78

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Deep Dual Relation Modeling for Egocentric Interaction Recognition

2019-05-31 · CVPR 2019 6 · Haoxin Li, Yijun Cai, Wei-Shi Zheng

Egocentric interaction recognition aims to recognize the camera wearer's interactions with the interactor who faces the camera wearer in egocentric videos. In such a human-human interaction analysis problem, it is crucia…

Relation

Attention-Oriented Action Recognition for Real-Time Human-Robot Interaction

2020-07-02 · Ziyang Song, Ziyi Yin, Zejian yuan, Chong Zhang 외

Despite the notable progress made in action recognition tasks, not much work has been done in action recognition specifically for human-robot interaction. In this paper, we deeply explore the characteristics of the actio…

Action RecognitionPose Estimation

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

2025-06-03 · Mengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang 외

Spatial reasoning is a key aspect of cognitive psychology and remains a major bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic s…

Object CountingSpatial Reasoning

A Neural Divide-and-Conquer Reasoning Framework for Image Retrieval from Linguistically Complex Text

2023-05-03 · Yunxin Li, Baotian Hu, Yuxin Ding, Lin Ma 외

Pretrained Vision-Language Models (VLMs) have achieved remarkable performance in image retrieval from text. However, their performance drops drastically when confronted with linguistically complex texts that they struggl…

Image RetrievalLogical ReasoningRetrieval

Seeing Beyond the Scene: Enhancing Vision-Language Models with Interactional Reasoning

2025-05-14 · Dayong Liang, Changmeng Zheng, Zhiyuan Wen, Yi Cai 외

Traditional scene graphs primarily focus on spatial relationships, limiting vision-language models' (VLMs) ability to reason about complex interactions in visual scenes. This paper addresses two key challenges: (1) conve…

Relation ExtractionScene Understanding