paper-with-me

홈 › Papers

LACE: Latent Visual Representation for Cross-Embodiment Learning

2026-05-16 · Yoo Sung Jang, Kanchana Ranasinghe, Cristina Mata, Yichi Zhang, Jorge Mendez-Mendez, Michael S. Ryoo arxiv

Cross-embodiment learning from human demonstrations is hindered by the visual gap between human and robot embodiments. While self-supervised learning (SSL) backbones encode rich inter-class semantics of general objects, we show they fail to establish correspondence between human and robot hands. We propose LACE, a framework that aligns human and robot visual representations in the latent space of these backbones by leveraging correspondences between shared body parts across embodiments as sparse supervision. These annotations can be automatically obtained via forward kinematics, and single robot demonstration is sufficient to train the model. Our semantic alignment loss matches distributions incurred by corresponding features, lifting patch-level supervision to semantic-level alignment, while a Gram loss preserves pretrained feature quality. This alignment enables robot policies to leverage abundant human data when robot demonstrations are scarce: in zero-shot transfer, policies using LACE-DINO outperform those using DINO by a large margin (65\%), with consistent gains in low-data regimes and out-of-distribution environments.

📄 PDF Abstract BibTeX arXiv:2605.16743

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

SCAR: Self-Supervised Continuous Action Representation Learning

2026-05-13 · Hongjia Liu, Fan Feng, Minghao Fu, Xinyue Wang 외 arxiv

Despite the central role of action in embodied intelligence, learning transferable action representations from visual transitions remains a fundamental challenge, particularly when world models must generalize across emb…

Representation Learning

Learning a Unified Latent Space for Cross-Embodiment Robot Control

2026-01-21 · Yashuai Yan, Dongheui Lee arxiv

We present a scalable framework for cross-embodiment humanoid robot control by learning a shared latent representation that unifies motion across humans and diverse humanoid platforms, including single-arm, dual-arm, and…

Contrastive Learning

Demo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment Imitation

2026-05-20 · Jingyang He, Guangrun Li, Jieyu Zhang, Chengkai Hou 외 arxiv

Robotic imitation learning is often treated as reproducing demonstrated actions, but actions are inherently embodiment-specific. When demonstrations come from humans or robots with different morphology, kinematics, or ac…

One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation

2026-03-15 · Juncheng Mu, Sizhe Yang, Hojin Bae, Feiyu Jia 외 arxiv

Cross-embodiment manipulation is crucial for enhancing the scalability of robot manipulation and reducing the high cost of data collection. However, the significant differences between embodiments, such as variations in …

Robot Manipulation

HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model

2026-05-29 · Xiang Zhu, Puzhen Yuan, Yichen Liu, Jianyu Chen arxiv

Learning generalizable vision-language-action (VLA) models from large-scale human videos is promising but challenging due to cross-embodiment discrepancies in both visual observations and executable actions. While latent…

Representation Learning