paper-with-me

홈 › Papers

EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning

2026-06-16 · Gaotian Wang, Kejia Ren, Andrew Morgan, Yiting Chen, Howard H. Qian, Podshara Chanrungmaneekul, Kaiyu Hang arxiv

Internet videos constitute the largest reservoir of embodied human manipulation knowledge, yet converting arbitrary RGB footage into actionable robot training data remains a major bottleneck. Existing lab- or factory-collected datasets are narrow in scale and diversity, limiting open-world robot learning. Instead of proposing a static dataset, we introduce EgoInfinity, a universal 4D hand-object interaction data engine that enables web-scale data generation for robot retargeting and learning. EgoInfinity is a modular engine integrating perception, segmentation, reconstruction, interaction-aware refinement, and retargeting to automate this traditionally unscalable video-to-action problem without human-in-the-loop annotation. Its modular design lets the engine continuously benefit from advances in any incorporated component. With EgoInfinity, in-the-wild human manipulation videos are lifted into agent-agnostic, metric 4D hand-object representations, including hand trajectories, 6-DoF object poses, and contact-relevant states. Rather than naively connecting standalone components, EgoInfinity combines cross-module metric calibration with interaction-aware refinement to improve physical reliability, reducing drift and contact inconsistencies common in pure visual reconstruction. We further propose a novel motion retargeter that compiles the recovered 3D hand motions into executable joint trajectories for diverse robot morphologies, enabling video-to-action retargeting on any robot from arbitrary viewpoints and shot sizes (e.g., the human body is only partially visible). We validate EgoInfinity across perception fidelity, kinematic feasibility, contact consistency, cross-embodiment generalization, and real-robot skill acquisition (e.g., grasping, cutting, wiping, and pouring), demonstrating a scalable bridge from internet videos to executable robot behavior for open-world robot learning.

📄 PDF Abstract BibTeX arXiv:2606.17385

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AffordPose: A Large-scale Dataset of Hand-Object Interactions with Affordance-driven Hand Pose

2023-09-16 · ICCV 2023 1 · Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu 외

How human interact with objects depends on the functional roles of the target objects, which introduces the problem of affordance-aware hand-object interaction. It requires a large number of human demonstrations for the …

DiversityObject

HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data Augmentation

2025-01-06 · Wentian Qu, Jiahe Li, Jian Cheng, Jian Shi 외

Understanding of bimanual hand-object interaction plays an important role in robotics and virtual reality. However, due to significant occlusions between hands and object as well as the high degree-of-freedom motions, it…

3DGSData Augmentationhand-object poseObject+1

Towards A Richer 2D Understanding of Hands at Scale

2023-09-21 · NeurIPS 2023 11

As humans, we learn a lot about how to interact with the world by observing others interacting with their hands. To help AI systems obtain a better understanding of hand interactions, we introduce a new model that produc…

DenseAttentionSeg: Segment Hands from Interacted Objects Using Depth Input

2019-03-29 · Zihao Bo, Hao Zhang, Junhai Yong, Feng Xu

We propose a real-time DNN-based technique to segment hand and object of interacting motions from depth inputs. Our model is called DenseAttentionSeg, which contains a dense attention mechanism to fuse information in dif…

Hand SegmentationObjectSegmentation

Affordance Diffusion: Synthesizing Hand-Object Interactions

2023-03-21 · CVPR 2023 1 · Yufei Ye, Xueting Li, Abhinav Gupta, Shalini De Mello 외

Recent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for synthesizing an entire image, texture trans…

DescriptiveImage GenerationObject