paper-with-me

홈 › Papers

Rethinking Visual Embodiment Dependence in Visuomotor Policies

2026-09-15 · Hongjie Fang, Yuxuan Lu, Chenxi Wang, Haoxiang Qin, Shirun Tang, Zihao He, Shangning Xia, Jingjing Chen, Wanxi Liu, Shiquan Wang, Cewu Lu arxiv

Visuomotor policies observe both the task scene and the acting embodiment, allowing embodiment-specific visual cues to influence action prediction. We study this phenomenon as visual embodiment dependence (VED) and show, through cue-conflict interventions across representative policies, that visible robot configuration can become a shortcut to task progress. Rather than eliminating VED, we argue that it should be structured around embodiment information that supports control and generalization. We realize this through embodiment canonicalization in 3D point clouds, replacing the original embodiment with a canonical end-effector representation (CER) that preserves control-relevant geometry while abstracting embodiment-specific morphology. Its editable form further enables configuration-decorrelation augmentation for unfamiliar robot configurations. Experiments show that embodiment canonicalization substantially improves human-to-robot policy transfer without robot demonstrations, while simply removing the embodiment is insufficient without preserving control-relevant geometry. We further find that CER itself can become a configuration shortcut when robot configuration becomes decoupled from task progress; configuration-decorrelation augmentation mitigates this failure mode and restores robust recovery without sacrificing performance on seen configurations. Together, these results show that robust visuomotor learning benefits from structuring, rather than removing, visual embodiment information. Project website: https://tonyfang.net/ved

📄 PDF Abstract BibTeX arXiv:2609.16815

Code (1)

DoHeyYeah/arXiV-favor ★ 1

Tasks

Point Clouds

Similar Papers 제목 키워드 기반

UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies

2025-10-02 · Harsh Gupta, Xiaofeng Guo, Huy Ha, Chuer Pan 외 arxiv

We introduce UMI-on-Air, a framework for embodiment-aware deployment of embodiment-agnostic manipulation policies. Our approach leverages diverse, unconstrained human demonstrations collected with a handheld gripper (UMI…

DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

2026-07-09 · Yunchao Yao, Zhuxiu Xu, Tianqi Zhang, Zixian Liu 외 arxiv

Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. …

Learning Adaptive Cross-Embodiment Visuomotor Policy with Contrastive Prompt Orchestration

2026-02-01 · Yuhang Zhang, Chao Yan, Jiaxi Yu, Jiaping Xiao 외 arxiv

Learning adaptive visuomotor policies for embodied agents remains a formidable challenge, particularly when facing cross-embodiment variations such as diverse sensor configurations and dynamic properties. Conventional le…

Contrastive Learning

Latent Policy Steering with Embodiment-Agnostic Pretrained World Models

2025-07-17 · Yiqi Wang, Mrinal Verghese, Jeff Schneider

Learning visuomotor policies via imitation has proven effective across a wide range of robotic domains. However, the performance of these policies is heavily dependent on the number of training demonstrations, which requ…

Maximizing Alignment with Minimal Feedback: Efficiently Learning Rewards for Visuomotor Robot Policy Alignment

2024-12-06 · Ran Tian, Yilin Wu, Chenfeng Xu, Masayoshi Tomizuka 외

Visuomotor robot policies, increasingly pre-trained on large-scale datasets, promise significant advancements across robotics domains. However, aligning these policies with end-user preferences remains a challenge, parti…