paper-with-me

Papers

Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning

2025-08-30 · Jiading Fang arxiv

This thesis introduces "Embodied Spatial Intelligence" to address the challenge of creating robots that can perceive and act in the real world based on natural language instructions. To bridge the gap between Large Language Models (LLMs) and physical embodiment, we present contributions on two fronts: scene representation and spatial reasoning. For perception, we develop robust, scalable, and accurate scene representations using implicit neural models, with contributions in self-supervised camera calibration, high-fidelity depth field generation, and large-scale reconstruction. For spatial reasoning, we enhance the spatial capabilities of LLMs by introducing a novel navigation benchmark, a method for grounding language in 3D, and a state-feedback mechanism to improve long-horizon decision-making. This work lays a foundation for robots that can robustly perceive their surroundings and intelligently act upon complex, language-based commands.

📄 PDF Abstract BibTeX arXiv:2509.00465

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Self in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied Intelligence

2026-07-14 · Zhishan Zou, Guoyan Sun, Zhiwei Wei, Jiancheng Pan 외 arxiv

Autonomous UAV systems increasingly rely on multimodal large language models (MLLMs) to operate in complex real-world environments. Such embodied scenarios require not only understanding the surrounding space but also ma…

EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks

2025-03-14 · Yi Zhang, Qiang Zhang, Xiaozhu Ju, Zhaoyang Liu 외

While multimodal large language models (MLLMs) have made groundbreaking progress in embodied intelligence, they still face significant challenges in spatial reasoning for complex long-horizon tasks. To address this gap, …

Spatial Reasoning

Reconstructing 4D Spatial Intelligence: A Survey

2025-07-28 · Yukang Cao, Jiahao Lu, Zhisheng Huang, Zhuowen Shen 외 arxiv

Reconstructing 4D spatial intelligence from visual observations has long been a central yet challenging task in computer vision, with broad real-world applications. These range from entertainment domains like movies, whe…

Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment

2025-08-29 · Jinzhou Tang, Jusheng zhang, Sidi Liu, Waikit Xiu 외 arxiv

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene u…

Scene UnderstandingQuestion Answering

SpatialPoint: Spatial-aware Point Prediction for Embodied Localization

2026-03-16 · Qiming Zhu, Zhirui Fang, Tianming Zhang, Chuanxiu Liu 외 arxiv

Embodied intelligence fundamentally requires a capability to determine where to act in 3D space. We formalize this requirement as embodied localization -- the problem of predicting executable 3D points conditioned on vis…

Spatial ReasoningRobot Navigation