paper-with-me

Spatial Reasoning

2개 벤치마크 · 논문 1,258편 · 이 태스크의 논문 보기 →

Benchmarks

6-DoF SpatialBench

결과 14개

EmbSpatial-Bench

결과 10개

Most implemented

Visual Instruction Tuning

2023-04-17 · 구현 13개

GPT-4 Technical Report

2023-03-15 · 구현 11개

Papers

EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning

2026-09-08 · Jingpu Yang, Fengxian Ji, Mingxuan Cui, Yilin Sun 외 arxiv

UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter t…

Video Question AnsweringSpatial Reasoning

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

2026-09-04 · Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan 외 arxiv

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offerin…

Spatial Reasoning

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

2026-09-03 · Yijun Yang, Shenghe Zheng, Wenbo Li, Jianhui Liu 외 hf

Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a …

Reinforcement LearningSpatial Reasoning

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

2026-09-01 · Ashwin Nedungadi, Stefan Oehmcke, Stefan Lüdtke hf

Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or si…

Spatial Reasoning

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

2026-08-31 · Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu 외 hf

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial pri…

Zero-shot GeneralizationReinforcement LearningInstruction FollowingSpatial Reasoning

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

2026-08-27 · Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui 외 hf

Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current M…

Spatial Reasoning

전체 1,258편 보기 →