paper-with-me

홈 › Papers

EgoDemoGen: Egocentric Demonstration Generation for Viewpoint Generalization in Robotic Manipulation

2025-09-26 · Yuan Xu, Jiabing Yang, Xiaofeng Wang, Yixiang Chen, Zheng Zhu, Bowen Fang, Guan Huang, Xinze Chen, Yun Ye, Qiang Zhang, Peiyan Li, Xiangnan Wu, Kai Wang, Bing Zhan, Shuo Lu, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang arxiv

Imitation learning based visuomotor policies have achieved strong performance in robotic manipulation, yet they often remain sensitive to egocentric viewpoint shifts. Unlike third-person viewpoint changes that only move the camera, egocentric shifts simultaneously alter both the camera pose and the robot action coordinate frame, making it necessary to jointly transfer action trajectories and synthesize corresponding observations under novel egocentric viewpoints. To address this challenge, we present EgoDemoGen, a framework that generates paired observation--action demonstrations under novel egocentric viewpoints through two key components: 1{)} EgoTrajTransfer, which transfers robot trajectories to the novel egocentric coordinate frame through motion-skill segmentation, geometry-aware transformation, and inverse kinematics filtering; and 2{)} EgoViewTransfer, a conditional video generation model that fuses a novel-viewpoint reprojected scene video and a robot motion video rendered from the transferred trajectory to synthesize photorealistic observations, trained with a self-supervised double reprojection strategy without requiring multi-viewpoint data. Experiments in simulation and real-world settings show that EgoDemoGen consistently improves policy success rates under both standard and novel egocentric viewpoints, with absolute gains of +24.6\% and +16.9\% in simulation and +16.0\% and +23.0\% on the real robot. Moreover, EgoViewTransfer achieves superior video generation quality for novel egocentric observations.

📄 PDF Abstract BibTeX arXiv:2509.22578

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

EgoAVFlow: Robot Policy Learning with Active Vision from Human Egocentric Videos via 3D Flow

2026-02-25 · Daesol Cho, Youngseok Jang, Danfei Xu, Sehoon Ha arxiv

Egocentric human videos provide a scalable source of manipulation demonstrations; however, deploying them on robots requires active viewpoint control to maintain task-critical visibility, which human viewpoint imitation …

EgoGuide: Egocentric Guidance for Efficient Robot-Free Demonstration Collection and Learning

2026-06-12 · Yue Xu, Mingtao Nie, Tianle Li, Hong Li 외 arxiv

Robot learning from real-world demonstrations is currently constrained by data scaling. Universal Manipulation Interface (UMI) provides an efficient robot-free data collection interface, yet current UMI-style pipelines o…

EgoMI: Learning Active Vision and Whole-Body Manipulation from Egocentric Human Demonstrations

2025-10-31 · Justin Yu, Yide Shentu, Di Wu, Pieter Abbeel 외 arxiv

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans act…

SID: Sliding into Distribution for Robust Few-Demonstration Manipulation

2026-05-13 · Yicheng Ma, Wei Yu, Zhian Su, Xidan Zhang 외 arxiv

Generalizing robotic manipulation across object poses, viewpoints, and dynamic disturbances is difficult, especially with only a few demonstrations. End-to-end visuomotor policies are expressive but data-hungry, while pl…

MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training

2025-09-26 · Haoyun Li, Ivan Zhang, Runqi Ouyang, Xiaofeng Wang 외 arxiv

Vision Language Action (VLA) models derive their generalization capability from diverse training data, yet collecting embodied robot interaction data remains prohibitively expensive. In contrast, human demonstration vide…

Pose Tracking