paper-with-me

Papers

ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos

2024-04-24 · Zerui Chen, ShiZhe Chen, Etienne Arlaud, Ivan Laptev, Cordelia Schmid

In this work, we aim to learn a unified vision-based policy for multi-fingered robot hands to manipulate a variety of objects in diverse poses. Though prior work has shown benefits of using human videos for policy learning, performance gains have been limited by the noise in estimated trajectories. Moreover, reliance on privileged object information such as ground-truth object states further limits the applicability in realistic scenarios. To address these limitations, we propose a new framework ViViDex to improve vision-based policy learning from human videos. It first uses reinforcement learning with trajectory guided rewards to train state-based policies for each video, obtaining both visually natural and physically plausible trajectories from the video. We then rollout successful episodes from state-based policies and train a unified visual policy without using any privileged information. We propose coordinate transformation to further enhance the visual point cloud representation, and compare behavior cloning and diffusion policy for the visual policy training. Experiments both in simulation and on the real robot demonstrate that ViViDex outperforms state-of-the-art approaches on three dexterous manipulation tasks.

📄 PDF Abstract BibTeX arXiv:2404.15709

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction

2026-09-07 · Ruoqu Chen, Feixiang Ruan, Liu Cao, Zihao Wang 외 arxiv

Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable vis…

Zero-shot Generalization

DexMV: Imitation Learning for Dexterous Manipulation from Human Videos

2021-08-12 · Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang 외

While significant progress has been made on understanding hand-object interactions in computer vision, it is still very challenging for robots to perform complex dexterous manipulation. In this paper, we propose a new pl…

Imitation Learningmotion retargetingTranslation

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

2026-06-20 · Yangtao Chen, Zixuan Chen, Peiyang Wang, Yong-Lu Li 외 arxiv

Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well aligned with r…

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos

2026-06-17 · Bhawna Paliwal, Haritheja Etukuru, William Liang, Pieter Abbeel 외 arxiv

How can we scalably generate data for robotic manipulation, especially on human-like platforms such as dexterous multi-fingered hands? Learning from human videos has recently emerged as a likely answer to this question. …

Robot Manipulation

Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction

2026-02-09 · Hongyi Chen, Tony Dong, Tiancheng Wu, Liquan Wang 외 arxiv

Multi-finger robotic hand manipulation and grasping are challenging due to the high-dimensional action space and the difficulty of acquiring large-scale training data. Existing approaches largely rely on human teleoperat…