paper-with-me

홈 › Papers

EgoMimic: Scaling Imitation Learning via Egocentric Video

2024-10-31 · Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, Danfei Xu

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through: (1) a system to capture human embodiment data using the ergonomic Project Aria glasses, (2) a low-cost bimanual manipulator that minimizes the kinematic gap to human data, (3) cross-domain data alignment techniques, and (4) an imitation learning architecture that co-trains on human and robot data. Compared to prior works that only extract high-level intent from human videos, our approach treats human and robot data equally as embodied demonstration data and learns a unified policy from both data sources. EgoMimic achieves significant improvement on a diverse set of long-horizon, single-arm and bimanual manipulation tasks over state-of-the-art imitation learning methods and enables generalization to entirely new scenes. Finally, we show a favorable scaling trend for EgoMimic, where adding 1 hour of additional hand data is significantly more valuable than 1 hour of additional robot data. Videos and additional information can be found at https://egomimic.github.io/

📄 PDF Abstract BibTeX arXiv:2410.24221

Code (1)

SimarKareer/EgoMimic 공식 구현 pytorch

Tasks

DiversityImitation Learning

Methods 이 논문이 사용한 방법론

ARiA 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Ego-Grounding for Personalized Question-Answering in Egocentric Videos

2026-04-02 · Junbin Xiao, Shenglang Zhang, Pengxiang Zhu, Angela Yao arxiv

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this …

HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

2026-06-18 · Juncheng Ma, Jianxin Bi, Yufan Deng, Xuanran Zhai 외 arxiv

Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to…

Object-Shot Enhanced Grounding Network for Egocentric Video

2025-05-07 · CVPR 2025 1 · Yisen Feng, Haoyu Zhang, Meng Liu, Weili Guan 외

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentr…

Video Grounding

EgoVSR: Towards High-Quality Egocentric Video Super-Resolution

2023-05-24 · Yichen Chi, Junhao Gu, Jiamiao Zhang, Wenming Yang 외

Due to the limitations of capture devices and scenarios, egocentric videos frequently have low visual quality, mainly caused by high compression and severe motion blur. With the increasing application of egocentric video…

Super-ResolutionVideo Super-Resolution

Opening the Vocabulary of Egocentric Actions

2023-08-22 · NeurIPS 2023 11 · Dibyadip Chatterjee, Fadime Sener, Shugao Ma, Angela Yao

Human actions in egocentric videos are often hand-object interactions composed from a verb (performed by the hand) applied to an object. Despite their extensive scaling up, egocentric datasets still face two limitations …

Action RecognitionObjectOpen Vocabulary Action Recognition