Spatial Reasoning
2개 벤치마크 · 논문 1,258편 · 이 태스크의 논문 보기 →
Benchmarks
6-DoF SpatialBench
EmbSpatial-Bench
Most implemented
Spatial Memory for Context Reasoning in Object Detection
Visual Instruction Tuning
GPT-4 Technical Report
Improved Baselines with Visual Instruction Tuning
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Long Range Arena: A Benchmark for Efficient Transformers
Papers
EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter t…
Video Question AnsweringSpatial ReasoningRoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offerin…
Spatial ReasoningUnfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a …
Reinforcement LearningSpatial ReasoningAutoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models
Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or si…
Spatial ReasoningLightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial pri…
Zero-shot GeneralizationReinforcement LearningInstruction FollowingSpatial ReasoningUrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current M…
Spatial Reasoning