paper-with-me

Papers

MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation

2026-02-06 · Haoming Wang, Qiyao Xue, Weichen Liu, Wei Gao arxiv

When embodied AI is expanding from traditional object detection and recognition to more advanced tasks of robot manipulation and actuation planning, visual spatial reasoning from the video inputs is necessary to perceive the spatial relationships of objects and guide device actions. However, existing visual language models (VLMs) have very weak capabilities in spatial reasoning due to the lack of knowledge about 3D spatial information, especially when the reasoning task involve complex spatial relations across multiple video frames. In this paper, we present a new inference-time computing technique for on-device embodied AI, namely \emph{MosaicThinker}, which enhances the on-device small VLM's spatial reasoning capabilities on difficult cross-frame reasoning tasks. Our basic idea is to integrate fragmented spatial information from multiple frames into a unified space representation of global semantic map, and further guide the VLM's spatial reasoning over the semantic map via a visual prompt. Experiment results show that our technique can greatly enhance the accuracy of cross-frame spatial reasoning on resource-constrained embodied AI devices, over reasoning tasks with diverse types and complexities.

📄 PDF Abstract BibTeX arXiv:2602.07082

Code (0)

등록된 구현이 없습니다.

Tasks

Robot ManipulationSpatial ReasoningObject Detection

Similar Papers 제목 키워드 기반

EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

2025-03-14 · Yi Zhang, Qiang Zhang, Xiaozhu Ju, Zhaoyang Liu 외

While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex long-horizon tasks. To address this gap, …

Spatial Reasoning

REM: Evaluating LLM Embodied Spatial Reasoning through Multi-Frame Trajectories

2025-11-30 · Jacob Thompson, Emiliano Garcia-Lopez, Yonatan Bisk arxiv

Humans build viewpoint-independent cognitive maps through navigation, enabling intuitive reasoning about object permanence and spatial relations. We argue that multimodal large language models (MLLMs), despite extensive …

Spatial Reasoning

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

2026-08-31 · Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu 외 hf

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial pri…

Zero-shot GeneralizationReinforcement LearningInstruction FollowingSpatial Reasoning

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

2026-05-13 · Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang 외 arxiv

In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to guide robot policies with explicit visu…

ERQA-Plus: A Diagnostic Benchmark for Reasoning in Embodied AI

2026-06-16 · Hong Yang, Basura Fernando arxiv

Generalist embodied agents require more than object recognition: they must reason about spatial relations, actions, procedures, human intentions, environmental constraints, and commonsense consequences from situated visu…

Question GenerationObject RecognitionQuestion AnsweringSpatial Reasoning