paper-with-me

홈 › Papers

Imitating What Works: Simulation-Filtered Modular Policy Learning from Human Videos

2026-02-13 · Albert J. Zhai, Kuo-Hao Zeng, Jiasen Lu, Ali Farhadi, Shenlong Wang, Wei-Chiu Ma arxiv

The ability to learn manipulation skills by watching videos of humans has the potential to unlock a new source of highly scalable data for robot learning. Here, we tackle prehensile manipulation, in which tasks involve grasping an object before performing various post-grasp motions. Human videos offer strong signals for learning the post-grasp motions, but they are less useful for learning the prerequisite grasping behaviors, especially for robots without human-like hands. A promising way forward is to use a modular policy design, leveraging a dedicated grasp generator to produce stable grasps. However, arbitrary stable grasps are often not task-compatible, hindering the robot's ability to perform the desired downstream motion. To address this challenge, we present Perceive-Simulate-Imitate (PSI), a framework for training a modular manipulation policy using human video motion data processed by paired grasp-trajectory filtering in simulation. This simulation step extends the trajectory data with grasp suitability labels, which allows for supervised learning of task-oriented grasping capabilities. We show through real-world experiments that our framework can be used to learn precise manipulation skills efficiently without any robot data, resulting in significantly more robust performance than using a grasp generator naively.

📄 PDF Abstract BibTeX arXiv:2602.13197

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning a Decision Module by Imitating Driver's Control Behaviors

2019-11-30 · Junning Huang, Sirui Xie, Jiankai Sun, Qiurui Ma 외

Autonomous driving systems have a pipeline of perception, decision, planning, and control. The decision module processes information from the perception module and directs the execution of downstream planning and control…

Autonomous DrivingImitation Learning

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

2025-07-01 · Shivansh Patel, Shraddhaa Mohan, Hanlin Mai, Unnat Jain 외

This work introduces Robots Imitating Generated Videos (RIGVid), a system that enables robots to perform complex manipulation tasks--such as pouring, wiping, and mixing--purely by imitating AI-generated videos, without r…

Point TrackingPose Tracking

Robust Asymmetric Learning in POMDPs

2020-12-31 · Andrew Warrington, J. Wilder Lavington, Adam Ścibior, Mark Schmidt 외

Policies for partially observed Markov decision processes can be efficiently learned by imitating policies for the corresponding fully observed Markov decision processes. Unfortunately, existing approaches for this kind …

Imitation Learning

Modular Primitives for High-Performance Differentiable Rendering

2020-11-06 · Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol 외

We present a modular differentiable renderer design that yields performance superior to previous methods by leveraging existing, highly optimized hardware graphics pipelines. Our design supports all crucial operations in…

AttributeInverse RenderingVocal Bursts Intensity Prediction

Perspective: Purposeful Failure in Artificial Life and Artificial Intelligence

2021-02-24 · Lana Sinapayen

Complex systems fail. I argue that failures can be a blueprint characterizing living organisms and biological intelligence, a control mechanism to increase complexity in evolutionary simulations, and an alternative to cl…

Artificial Life