paper-with-me

Papers

Latent Policy Steering with Embodiment-Agnostic Pretrained World Models

2025-07-17 · Yiqi Wang, Mrinal Verghese, Jeff Schneider

Learning visuomotor policies via imitation has proven effective across a wide range of robotic domains. However, the performance of these policies is heavily dependent on the number of training demonstrations, which requires expensive data collection in the real world. In this work, we aim to reduce data collection efforts when learning visuomotor robot policies by leveraging existing or cost-effective data from a wide range of embodiments, such as public robot datasets and the datasets of humans playing with objects (human data from play). Our approach leverages two key insights. First, we use optic flow as an embodiment-agnostic action representation to train a World Model (WM) across multi-embodiment datasets, and finetune it on a small amount of robot data from the target embodiment. Second, we develop a method, Latent Policy Steering (LPS), to improve the output of a behavior-cloned policy by searching in the latent space of the WM for better action sequences. In real world experiments, we observe significant improvements in the performance of policies trained with a small amount of data (over 50% relative improvement with 30 demonstrations and over 20% relative improvement with 50 demonstrations) by combining the policy with a WM pretrained on two thousand episodes sampled from the existing Open X-embodiment dataset across different robots or a cost-effective human dataset from play.

📄 PDF Abstract BibTeX arXiv:2507.13340

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments

2026-06-01 · Kihyun Kim, Chaeyun Kim, Jongho Shin, Taeyoun Kwon 외 arxiv

Learning a good action embedding space is fundamental to scalable robot policy learning, yet existing methods treat action latents as task-specific intermediates rather than first-class representations. The resulting lat…

EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment

2026-06-11 · Shihefeng Wang, Kangchen Lv, Mingrui Yu, Xiang Li arxiv

Scalable robot imitation learning relies on large-scale heterogeneous data from diverse robots or body-free data, making Cartesian end-effector actions a key interface for embodiment-agnostic policy learning. However, en…

Collision Avoidance

KITE: Decoupling Kinematics and Interaction for Zero-Shot Cross-Embodiment Manipulation

2026-06-20 · Qianxu Wang, Kuan Fang arxiv

Generalizing manipulation policies across robot embodiments remains difficult because standard policies entangle task reasoning with embodiment-specific motor control. We study zero-shot cross-embodiment manipulation, wh…

TTT-VLA: Test-Time Latent Prompt Optimization for Vision-Language-Action Models

2026-06-02 · Wenbo Zhang, Jianxiong Li, Shuai Yang, Sijin Chen 외 arxiv

Vision-Language-Action (VLA) models trained on large-scale data have made remarkable progress, but they remain vulnerable to distribution shifts at deployment time. Recent VLA models suggest that prompts can serve as an …

Learning a Unified Latent Space for Cross-Embodiment Robot Control

2026-01-21 · Yashuai Yan, Dongheui Lee arxiv

We present a scalable framework for cross-embodiment humanoid robot control by learning a shared latent representation that unifies motion across humans and diverse humanoid platforms, including single-arm, dual-arm, and…

Contrastive Learning