paper-with-me

홈 › Papers

Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

2024-06-24 · Junbang Liang, Ruoshi Liu, Ege Ozguroglu, Sruthi Sudhakar, Achal Dave, Pavel Tokmakov, Shuran Song, Carl Vondrick

A key challenge in manipulation is learning a policy that can robustly generalize to diverse visual environments. A promising mechanism for learning robust policies is to leverage video generative models, which are pretrained on large-scale datasets of internet videos. In this paper, we propose a visuomotor policy learning framework that fine-tunes a video diffusion model on human demonstrations of a given task. At test time, we generate an example of an execution of the task conditioned on images of a novel scene, and use this synthesized execution directly to control the robot. Our key insight is that using common tools allows us to effortlessly bridge the embodiment gap between the human hand and the robot manipulator. We evaluate our approach on four tasks of increasing complexity and demonstrate that harnessing internet-scale generative models allows the learned policy to achieve a significantly higher degree of generalization than existing behavior cloning approaches.

📄 PDF Abstract BibTeX arXiv:2406.16862

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Learning Navigation Subroutines from Egocentric Videos

2019-05-29 · Ashish Kumar, Saurabh Gupta, Jitendra Malik

Planning at a higher level of abstraction instead of low level torques improves the sample efficiency in reinforcement learning, and computational efficiency in classical planning. We propose a method to learn such hiera…

Computational EfficiencyPseudo LabelReinforcement Learning

Learning to Drive by Watching YouTube Videos: Action-Conditioned Contrastive Policy Pretraining

2022-04-05 · Qihang Zhang, Zhenghao Peng, Bolei Zhou

Deep visuomotor policy learning, which aims to map raw visual observation to action, achieves promising results in control tasks such as robotic manipulation and autonomous driving. However, it requires a huge number of …

Autonomous DrivingImitation Learning

Out-of-Distribution Recovery with Object-Centric Keypoint Inverse Policy for Visuomotor Imitation Learning

2024-11-05 · George Jiayuan Gao, Tianyu Li, Nadia Figueroa

We propose an object-centric recovery (OCR) framework to address the challenges of out-of-distribution (OOD) scenarios in visuomotor policy learning. Previous behavior cloning (BC) methods rely heavily on a large amount …

Continual LearningImitation LearningObjectOptical Character Recognition (OCR)

SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment

2025-11-11 · Rong Xue, Jiageng Mao, Mingtong Zhang, Yue Wang arxiv

Developing efficient and accurate visuomotor policies poses a central challenge in robotic imitation learning. While recent rectified flow approaches have advanced visuomotor policy learning, they suffer from a key limit…

Spatial Policy: Guiding Visuomotor Robotic Manipulation with Spatial-Aware Modeling and Reasoning

2025-08-21 · Yijun Liu, Yuwei Liu, Yuan Meng, Jieheng Zhang 외 arxiv

Vision-centric hierarchical embodied models have demonstrated strong potential. However, existing methods lack spatial awareness capabilities, limiting their effectiveness in bridging visual plans to actionable control i…

Spatial ReasoningVideo Generation