paper-with-me

홈 › Papers

Hand-Object Interaction Pretraining from Videos

2024-09-12 · Himanshu Gaurav Singh, Antonio Loquercio, Carmelo Sferrazza, Jane Wu, Haozhi Qi, Pieter Abbeel, Jitendra Malik

We present an approach to learn general robot manipulation priors from 3D hand-object interaction trajectories. We build a framework to use in-the-wild videos to generate sensorimotor robot trajectories. We do so by lifting both the human hand and the manipulated object in a shared 3D space and retargeting human motions to robot actions. Generative modeling on this data gives us a task-agnostic base policy. This policy captures a general yet flexible manipulation prior. We empirically demonstrate that finetuning this policy, with both reinforcement learning (RL) and behavior cloning (BC), enables sample-efficient adaptation to downstream tasks and simultaneously improves robustness and generalizability compared to prior approaches. Qualitative experiments are available at: \url{https://hgaurav2k.github.io/hop/}.

📄 PDF Abstract BibTeX arXiv:2409.08273

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectReinforcement Learning (RL)Robot Manipulation

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos

2026-02-10 · Juncheng Mu, Sizhe Yang, Yiming Bao, Hojin Bae 외 arxiv

Data scarcity fundamentally limits the generalization of bimanual dexterous manipulation, as real-world data collection for dexterous hands is expensive and labor-intensive. Human manipulation videos, as a direct carrier…

Data AugmentationVideo Generation

Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos

2025-10-24 · Qixiu Li, Yu Deng, Yaobo Liang, Lin Luo 외 arxiv

This paper presents a novel approach for pretraining robotic manipulation Vision-Language-Action (VLA) models using a large corpus of unscripted real-life video recordings of human hand activities. Treating human hand as…

Learning to Imitate Object Interactions from Internet Videos

2022-11-23 · Austin Patel, Andrew Wang, Ilija Radosavovic, Jitendra Malik

We study the problem of imitating object interactions from Internet videos. This requires understanding the hand-object interactions in 4D, spatially in 3D and over time, which is challenging due to mutual hand-object oc…

Object

ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos

2026-02-05 · Yuantao Chen, Jiahao Chang, Chongjie Ye, Chaoran Zhang 외 arxiv

The ubiquity of monocular videos capturing daily hand-object interactions presents a valuable resource for embodied intelligence. While 3D hand reconstruction from in-the-wild videos has seen significant progress, recons…

3D Object Reconstruction

WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos

2026-02-25 · Yufei Ye, Jiaman Li, Ryan Rong, C. Karen Liu arxiv

Egocentric manipulation videos are highly challenging due to severe occlusions during interactions and frequent object entries and exits from the camera view as the person moves. Current methods typically focus on recove…

Pose Estimation