paper-with-me

Papers

A Persistent Spatial Semantic Representation for High-level Natural Language Instruction Execution

2021-07-12 · Valts Blukis, Chris Paxton, Dieter Fox, Animesh Garg, Yoav Artzi

Natural language provides an accessible and expressive interface to specify long-term tasks for robotic agents. However, non-experts are likely to specify such tasks with high-level instructions, which abstract over specific robot actions through several layers of abstraction. We propose that key to bridging this gap between language and robot actions over long execution horizons are persistent representations. We propose a persistent spatial semantic representation method, and show how it enables building an agent that performs hierarchical reasoning to effectively execute long-term tasks. We evaluate our approach on the ALFRED benchmark and achieve state-of-the-art results, despite completely avoiding the commonly used step-by-step instructions.

📄 PDF Abstract BibTeX arXiv:2107.05612

Code (1)

valtsblukis/hlsm pytorch

Similar Papers 제목 키워드 기반

OpenSPM: An Environment-Transferable Robotic Key Spatial Pose Memory and Closed-Loop High-Frequency Flow-Matching Action Generation Model

2026-06-29 · Iok Tong Lei, Qingchen Xie, Yifan Wang, Yap Ying Jie 외 arxiv

Open-environment tabletop robotic manipulation requires systems to possess semantic understanding, precise geometric pose estimation, and high-frequency action generation. While end-to-end vision-language-action (VLA) mo…

Pose Estimation

Vision-Language Memory for Spatial Reasoning

2025-11-25 · Zuntao Liu, Yi Du, Taimeng Fu, Shaoshu Su 외 arxiv

Spatial reasoning is a critical capability for intelligent robots, yet current vision-language models (VLMs) still fall short of human-level performance in video-based spatial reasoning. This gap mainly stems from two ch…

Spatial Reasoning

GSMem: 3D Gaussian Splatting as Persistent Spatial Memory for Zero-Shot Embodied Exploration and Reasoning

2026-03-19 · Yiren Lu, Yi Du, Disheng Liu, Yunlai Zhou 외 arxiv

Effective embodied exploration requires agents to accumulate and retain spatial knowledge over time. However, existing scene representations, such as discrete scene graphs or static view-based snapshots, lack \textit{pos…

Question Answering

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation

2026-07-14 · Mohan Liu, Zhihao Gu, Xuanyu Chen, Haitian Zhang 외 arxiv

Vision-Language-Action (VLA) models have emerged as a powerful end-to-end paradigm for robotic manipulation by mapping language instructions and 2D visual inputs directly to actions. However, these models lack an explici…

Point Clouds

R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space

2025-12-17 · Tin Stribor Sohn, Maximilian Dillitzer, Jason J. Corso, Eric Sax arxiv

Humans perceive and reason about their surroundings in four dimensions by building persistent, structured internal representations that encode semantic meaning, spatial layout, and temporal dynamics. These multimodal mem…

Natural Language QueriesQuestion Answering