paper-with-me

홈 › Papers

Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing

2026-05-05 · Zhiyuan Li, Wenyan Yang, Wenshuai Zhao, Yue Ma, Yuanpeng Tu, Pekka Marttinen, Joni Pajarinen arxiv

Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled representations, where task-relevant information is coupled with human-specific kinematics, limiting their adaptability. We propose a generative framework for cross-embodiment video editing that directly addresses this by learning explicitly disentangled task and embodiment representations. Our method factorizes a demonstration video into two orthogonal latent spaces by enforcing a dual contrastive objective: it minimizes mutual information between the spaces to ensure independence while maximizing intra-space consistency to create stable representations. A parameter-efficient adapter injects these latent codes into a frozen video diffusion model, enabling the synthesis of a coherent robot execution video from a single human demonstration, without requiring paired cross-embodiment data. Experiments show our approach generates temporally consistent and morphologically accurate robot demonstrations, offering a scalable solution to leverage internet-scale human video for robot learning.

📄 PDF Abstract BibTeX arXiv:2605.03637

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation

2026-05-12 · Yiren Song, Xiyao Deng, Pei Yang, Yihan Wang 외 arxiv

Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge …

Video Generation

BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks

2026-02-03 · Yixiang Chen, Peiyan Li, Jiabing Yang, Keji He 외 arxiv

Embodied world models have emerged as a promising paradigm in robotics, most of which leverage large-scale Internet videos or pretrained video generation models to enrich visual and motion priors. However, they still fac…

Video Generation

Motion-Focused Latent Action Enables Cross-Embodiment VLA Training from Human EgoVideos

2026-06-17 · Runze Xu, Yiluo Zhang, Jian Wang, Yu Wang 외 arxiv

Training generalist Vision-Language-Action(VLA) models typically requires massive, diverse robotic datasets with high-fidelity action annotations. While egocentric human manipulation videos are abundant and capture signi…

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

2026-08-13 · Dingyi Rong, Yue Shi, Chaofan Ma, Jiezhang Cao 외 arxiv

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behaviora…

Robot ManipulationVideo Generation

LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment

2026-04-12 · Yifu Xu, Bokai Lin, Xinyu Zhan, Hongjie Fang 외 arxiv

Scaling up robot learning is hindered by the scarcity of robotic demonstrations, whereas human videos offer a vast, untapped source of interaction data. However, bridging the embodiment gap between human hands and robot …