paper-with-me

Papers

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding

2025-11-21 · Nikolay Nikolov, Giuliano Albanese, Sombit Dey, Aleksandar Yanev, Luc Van Gool, Jan-Nico Zaech, Danda Pani Paudel arxiv

Robotic Foundation Models (RFMs) hold great promise as generalist, end-to-end systems for robot control. Yet their ability to generalize across new environments, tasks, and embodiments remains limited. We argue that a major bottleneck lies in their foundations: most RFMs are built by fine-tuning internet-pretrained Vision-Language Models (VLMs). However, these VLMs are trained on 2D image-language tasks and lack the 3D spatial reasoning inherently required for embodied control in the 3D world. Bridging this gap directly with large-scale robotic data is costly and difficult to scale. Instead, we propose to enrich easy-to-collect non-robotic image data with 3D annotations and enhance a pretrained VLM with 3D understanding capabilities. Following this strategy, we train SPEAR-VLM, a 3D-aware VLM that infers object coordinates in 3D space from a single 2D image. Building on SPEAR-VLM, we introduce our main contribution, $~\textbf{SPEAR-1}$: a robotic foundation model that integrates grounded 3D perception with language-instructed embodied control. Trained on $\sim$45M frames from 24 Open X-Embodiment datasets, SPEAR-1 outperforms or matches state-of-the-art models such as $π_0$-FAST and $π_{0.5}$, while it uses 20$\times$ fewer robot demonstrations. This carefully-engineered training strategy unlocks new VLM capabilities and as a consequence boosts the reliability of embodied control beyond what is achievable with only robotic data. We make our model weights and 3D-annotated datasets publicly available at https://spear.insait.ai.

📄 PDF Abstract BibTeX arXiv:2511.17411

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Generalist Robot Manipulation beyond Action Labeled Data

2025-09-24 · Alexander Spiridonov, Jan-Nico Zaech, Nikolay Nikolov, Luc Van Gool 외 arxiv

Recent advances in generalist robot manipulation leverage pre-trained Vision-Language Models (VLMs) and large-scale robot demonstrations to tackle diverse tasks in a zero-shot manner. A key challenge remains: scaling hig…

Robot ManipulationPoint Clouds

Beyond Viewpoint Generalization: What Multi-View Demonstrations Offer and How to Synthesize Them for Robot Manipulation?

2026-03-23 · Boyang Cai, Qiwei Liang, Jiawei Li, Shihang Weng 외 arxiv

Does multi-view demonstration truly improve robot manipulation, or merely enhance cross-view robustness? We present a systematic study quantifying the performance gains, scaling behavior, and underlying mechanisms of mul…

Robot Manipulation

Habitat-Web: Learning Embodied Object-Search Strategies from Human Demonstrations at Scale

2022-04-07 · CVPR 2022 1 · Ram Ramrakhya, Eric Undersander, Dhruv Batra, Abhishek Das

We present a large-scale study of imitating human demonstrations on tasks that require a virtual robot to search for objects in new environments -- (1) ObjectGoal Navigation (e.g. 'find & go to a chair') and (2) Pick&Pla…

Imitation LearningObjectGoal NavigationReinforcement Learning (RL)

MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

2023-10-26 · Ajay Mandlekar, Soroush Nasiriany, Bowen Wen, Iretiayo Akinola 외

Imitation learning from a large set of human demonstrations has proved to be an effective paradigm for building capable robot agents. However, the demonstrations can be extremely costly and time-consuming to collect. We …

Imitation Learning

BC-Z: Zero-Shot Task Generalization with Robotic Imitation Learning

2022-02-04 · Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler 외

In this paper, we study the problem of enabling a vision-based robotic manipulation system to generalize to novel tasks, a long-standing challenge in robot learning. We approach the challenge from an imitation learning p…

Imitation Learning