paper-with-me

홈 › Papers

Verifiably Following Complex Robot Instructions with Foundation Models

2024-02-18 · Benedict Quartey, Eric Rosen, Stefanie Tellex, George Konidaris

When instructing robots, users want to flexibly express constraints, refer to arbitrary landmarks, and verify robot behavior, while robots must disambiguate instructions into specifications and ground instruction referents in the real world. To address this problem, we propose Language Instruction grounding for Motion Planning (LIMP), an approach that enables robots to verifiably follow complex, open-ended instructions in real-world environments without prebuilt semantic maps. LIMP constructs a symbolic instruction representation that reveals the robot's alignment with an instructor's intended motives and affords the synthesis of correct-by-construction robot behaviors. We conduct a large-scale evaluation of LIMP on 150 instructions across five real-world environments, demonstrating its versatility and ease of deployment in diverse, unstructured domains. LIMP performs comparably to state-of-the-art baselines on standard open-vocabulary tasks and additionally achieves a 79\% success rate on complex spatiotemporal instructions, significantly outperforming baselines that only reach 38\%. See supplementary materials and demo videos at https://robotlimp.github.io

📄 PDF Abstract BibTeX arXiv:2402.11498

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Planning

Similar Papers 제목 키워드 기반

Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation

2024-11-12 · Sonia Raychaudhuri, Duy Ta, Katrina Ashton, Angel X. Chang 외

Large scale scenes such as multifloor homes can be robustly and efficiently mapped with a 3D graph of landmarks estimated jointly with robot poses in a factor graph, a technique commonly used in commercial robots such as…

Instruction FollowingObjectVision-Language Navigation

AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents

2024-01-23 · Michael Ahn, Debidatta Dwibedi, Chelsea Finn, Montse Gonzalez Arenas 외

Foundation models that incorporate language, vision, and more recently actions have revolutionized the ability to harness internet scale data to reason about useful tasks. However, one of the key challenges of training e…

Instruction FollowingScene Understanding

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

2026-07-16 · Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li 외 hf

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-th…

CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

2025-08-19 · Catherine Glossop, William Chen, Arjun Bhorkar, Dhruv Shah 외 arxiv

Generalist robots should be able to understand and follow user instructions. Despite providing a powerful architecture for mapping open-vocabulary language instructions to robot actions, current vision-language-action (V…

Vision-Language NavigationInstruction Following

Complex Logical Instruction Generation

2025-08-12 · Mian Zhang, Shujian Liu, Sixun Dong, Ming Yin 외 arxiv

Instruction following has catalyzed the recent era of Large Language Models (LLMs) and is the foundational skill underpinning more advanced capabilities such as reasoning and agentic behaviors. As tasks grow more challen…

Instruction Following