paper-with-me

Papers Spatial Reasoning

“Spatial Reasoning” 태그가 달린 논문 1,258편 · 필터 해제

EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning

2026-09-08 · Jingpu Yang, Fengxian Ji, Mingxuan Cui, Yilin Sun 외 arxiv

UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter t…

Video Question AnsweringSpatial Reasoning

RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

2026-09-04 · Zhenxuan Fan, Bo Zhang, Yutong Lin, Yuqian Yuan 외 arxiv

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offerin…

Spatial Reasoning

Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning

2026-09-03 · Yijun Yang, Shenghe Zheng, Wenbo Li, Jianhui Liu 외 hf

Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a …

Reinforcement LearningSpatial Reasoning

Autoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models

2026-09-01 · Ashwin Nedungadi, Stefan Oehmcke, Stefan Lüdtke hf

Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or si…

Spatial Reasoning

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

2026-08-31 · Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu 외 hf

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial pri…

Zero-shot GeneralizationReinforcement LearningInstruction FollowingSpatial Reasoning

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

2026-08-27 · Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui 외 hf

Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current M…

Spatial Reasoning

Comparative Evaluation of 3D Reconstruction Methods for Immersive Visualization of Laboratory Objects

2026-08-27 · Brian De La Cruz, Aaron Y. Zhao, Maitrey Gramopadhye, Sawyer J. Lazar 외 arxiv

In this study, we examined whether current 3D reconstruction methods can support the creation of realistic holographic representations of laboratory objects for educational use. In this regard, we compared four approache…

Spatial Reasoning3D Reconstruction

Skill Issue: Are Skills Language-Invariant in LLMs?

2026-08-26 · Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett 외 arxiv

Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsiste…

Spatial Reasoning

Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?

2026-08-24 · Xiwei Liu, Yulong Li, Xinlin Zhuang, Xuhui Li 외 arxiv

Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires …

Object LocalizationSpatial Reasoning

InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation

2026-08-24 · Mengao Zhao, Ziang Li, Chaodong Huang, Mengchen Ma 외 arxiv

Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies ac…

Instruction FollowingRobot ManipulationSpatial Reasoning

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

2026-08-24 · Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan 외 arxiv

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects i…

Novel View SynthesisSpatial Reasoning

GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding

2026-08-22 · Md Abrar Jahin, Md Rizwan Parvez hf

Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GU…

Spatial Reasoning

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

2026-08-21 · Lars Benedikt Kaesberg, Tianyu Yang, Florian Valentin Wunderlich, Terry Ruas 외 arxiv

Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less cle…

Multimodal ReasoningSpatial ReasoningVisual GroundingObject Detection

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans

2026-08-21 · Simon Vincent Abel, Heiko Hillenhagen, Michael Götz, Timo Ropinski 외 arxiv

Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding. While modern vision-language m…

Spatial Reasoning

Bridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation

2026-08-21 · Derui Li, Qian Qiao, Yuhao Sun, Wenhao Guo 외 arxiv

Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered…

Spatial ReasoningImage Generation

Projector Is All You Train

2026-08-20 · Nyx Iskandar, Saathvik Selvan, Slater Victoroff arxiv

The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuni…

Spatial Reasoning3D Classification

Keep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification

2026-08-20 · Fuad Hasan, Chul Min Yeum arxiv

Disaster damage is spatial: buildings rarely fail in isolation. Yet using spatial context for damage classification remains surprisingly underexplored, and many pipelines still rely primarily on per-building appearance c…

Spatial Reasoning

GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery

2026-08-19 · Chaowei Wang, Yan Di, Jingjun Sun, Baozhe Liu 외 arxiv

Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a natural language description. However, the abundance of small, densely distribut…

Spatial ReasoningVisual Grounding

aDSL: Agentic 3D Creation via Joint Agent-Program Design

2026-08-18 · Rui-Huan Wang, Si-Tong Wei, Jia-Qi He, Heng-Yi Wei 외 arxiv

Programmatic representations provide a compelling paradigm for 3D content creation, enabling fine-grained edits, interpretability, and explicit structural control. Yet, agentic workflows that rely on large language model…

Spatial Reasoning

Spatial Message Passing in Language Space for Pathology Image Interpretation

2026-08-14 · Jing-Cheng Yang, Hao-Jung Wang, Jinhao Du, Yang Hu 외 arxiv

Multimodal Large Language Models (MLLMs) can generate pathological descriptions from histological images, but gigapixel Whole Slide Images (WSIs) exceed their visual context limits. The standard tiling workaround makes W…

Spatial Reasoning
1–20 / 1,258 다음 →