Papers Spatial Reasoning
“Spatial Reasoning” 태그가 달린 논문 1,258편 · 필터 해제
EgoSIS: From Factorized Visual Ego-Transitions to Motion-Canonical Spatial Evidence for UAV Reasoning
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter t…
Video Question AnsweringSpatial ReasoningRoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?
Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offerin…
Spatial ReasoningUnfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a …
Reinforcement LearningSpatial ReasoningAutoregressive Mosaics: Probing 2D Spatial Reasoning in Text-Only Language Models
Large language models (LLMs) trained only on text and code can sometimes generate programs that draw recognizable images. However, it is unclear whether this reflects an internal representation of 2D spatial layout or si…
Spatial ReasoningLightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial pri…
Zero-shot GeneralizationReinforcement LearningInstruction FollowingSpatial ReasoningUrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current M…
Spatial ReasoningComparative Evaluation of 3D Reconstruction Methods for Immersive Visualization of Laboratory Objects
In this study, we examined whether current 3D reconstruction methods can support the creation of realistic holographic representations of laboratory objects for educational use. In this regard, we compared four approache…
Spatial Reasoning3D ReconstructionSkill Issue: Are Skills Language-Invariant in LLMs?
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsiste…
Spatial ReasoningGrounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?
Vision-language models (VLMs) can answer spatial questions, yet the mechanisms connecting object grounding to spatial reasoning remain poorly understood. It is underexplored whether spatial reasoning internally requires …
Object LocalizationSpatial ReasoningInstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation
Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies ac…
Instruction FollowingRobot ManipulationSpatial ReasoningObject-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation
Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects i…
Novel View SynthesisSpatial ReasoningGUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GU…
Spatial ReasoningIs Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an interaction between visual grounding and downstream reasoning. What remains less cle…
Multimodal ReasoningSpatial ReasoningVisual GroundingObject DetectionA Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding. While modern vision-language m…
Spatial ReasoningBridging Language and Spherical Space: Object-Centric Control for Text-to-Panorama Generation
Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered…
Spatial ReasoningImage GenerationProjector Is All You Train
The typical training process of a multimodal large language model (MLLM) involves adapting both the language model backbone and the projector between the backbone and a modality-specific encoder. We ask whether fine-tuni…
Spatial Reasoning3D ClassificationKeep Your Friends Close, and the Right Neighbours Closer: Disaster-Conditioned Kernel-Regularized Graph Attention for Building Damage Classification
Disaster damage is spatial: buildings rarely fail in isolation. Yet using spatial context for damage classification remains surprisingly underexplored, and many pipelines still rely primarily on per-building appearance c…
Spatial ReasoningGrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery
Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a natural language description. However, the abundance of small, densely distribut…
Spatial ReasoningVisual GroundingaDSL: Agentic 3D Creation via Joint Agent-Program Design
Programmatic representations provide a compelling paradigm for 3D content creation, enabling fine-grained edits, interpretability, and explicit structural control. Yet, agentic workflows that rely on large language model…
Spatial ReasoningSpatial Message Passing in Language Space for Pathology Image Interpretation
Multimodal Large Language Models (MLLMs) can generate pathological descriptions from histological images, but gigapixel Whole Slide Images (WSIs) exceed their visual context limits. The standard tiling workaround makes W…
Spatial Reasoning