paper-with-me

홈 › Papers

Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos

2025-02-14 · Weirui Ye, Fangchen Liu, Zheng Ding, Yang Gao, Oleh Rybkin, Pieter Abbeel

Simulation offers a promising approach for cheaply scaling training data for generalist policies. To scalably generate data from diverse and realistic tasks, existing algorithms either rely on large language models (LLMs) that may hallucinate tasks not interesting for robotics; or digital twins, which require careful real-to-sim alignment and are hard to scale. To address these challenges, we introduce Video2Policy, a novel framework that leverages internet RGB videos to reconstruct tasks based on everyday human behavior. Our approach comprises two phases: (1) task generation in simulation from videos; and (2) reinforcement learning utilizing in-context LLM-generated reward functions iteratively. We demonstrate the efficacy of Video2Policy by reconstructing over 100 videos from the Something-Something-v2 (SSv2) dataset, which depicts diverse and complex human behaviors on 9 different tasks. Our method can successfully train RL policies on such tasks, including complex and challenging tasks such as throwing. Finally, we show that the generated simulation data can be scaled up for training a general policy, and it can be transferred back to the real robot in a Real2Sim2Real way.

📄 PDF Abstract BibTeX arXiv:2502.09886

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos

2025-04-08 · Alexey Gavryushin, Xi Wang, Robert J. S. Malate, Chenyu Yang 외

Large-scale egocentric video datasets capture diverse human activities across a wide range of scenarios, offering rich and detailed insights into how humans interact with objects, especially those that require fine-grain…

Sim-and-Real Co-Training: A Simple Recipe for Vision-Based Robotic Manipulation

2025-03-31 · Abhiram Maddukuri, Zhenyu Jiang, Lawrence Yunliang Chen, Soroush Nasiriany 외

Large real-world robot datasets hold great potential to train generalist robot models, but scaling real-world human data collection is time-consuming and resource-intensive. Simulation has great potential in supplementin…

Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction

2026-09-07 · Ruoqu Chen, Feixiang Ruan, Liu Cao, Zihao Wang 외 arxiv

Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable vis…

Zero-shot Generalization

Imitating What Works: Simulation-Filtered Modular Policy Learning from Human Videos

2026-02-13 · Albert J. Zhai, Kuo-Hao Zeng, Jiasen Lu, Ali Farhadi 외 arxiv

The ability to learn manipulation skills by watching videos of humans has the potential to unlock a new source of highly scalable data for robot learning. Here, we tackle prehensile manipulation, in which tasks involve g…

ManiBox: Enhancing Spatial Grasping Generalization via Scalable Simulation Data Generation

2024-11-04 · Hengkai Tan, Xuezhou Xu, Chengyang Ying, Xinyi Mao 외

Learning a precise robotic grasping policy is crucial for embodied agents operating in complex real-world manipulation tasks. Despite significant advancements, most models still struggle with accurate spatial positioning…

Robotic Grasping