paper-with-me

홈 › Papers

RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience

2026-08-19 · Yaowei Guo, Zeng Tao, Yuxin Jiang, Yunuo Chen, Zhiyang Dou, Yuxiang Ma, Yin Yang, Demetri Terzopoulos, Ying Jiang, Chenfanfu Jiang arxiv

Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning. Project webpage: https://roboedit.github.io/

📄 PDF Abstract BibTeX arXiv:2608.18948

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WholeBodyVLA: Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control

2025-12-11 · Haoran Jiang, Jin Chen, Qingwen Bu, Li Chen 외 arxiv

Humanoid robots require precise locomotion and dexterous manipulation to perform challenging loco-manipulation tasks. Yet existing approaches, modular or end-to-end, are deficient in manipulation-aware locomotion. This c…

Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

2026-05-14 · Yubo Zhao, Yujin Chai, Yunao Dong, Chengfeng Zhao 외 arxiv

Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can reconstruct temporally coherent human an…

Robot Self-Improvement via Human-Video Dynamics Models

2026-06-19 · Hanzhi Chen, Anran Zhang, Simon Schaefer, Kejia Chen 외 arxiv

A central question in robot learning is how to acquire skills from the kinds of data that humans learn from: passive observation, embodied practice, and the experience of failure. Human videos provide the first of these …

Cross-Domain Transfer via Semantic Skill Imitation

2022-12-14 · Karl Pertsch, Ruta Desai, Vikash Kumar, Franziska Meier 외

We propose an approach for semantic imitation, which uses demonstrations from a source domain, e.g. human videos, to accelerate reinforcement learning (RL) in a different target domain, e.g. a robotic manipulator in a si…

Reinforcement Learning (RL)Robot Manipulation

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos

2026-06-17 · Bhawna Paliwal, Haritheja Etukuru, William Liang, Pieter Abbeel 외 arxiv

How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. …

Robot Manipulation