paper-with-me

Papers

Egocentric Object Manipulation Graphs

2020-06-05 · Eadom Dessalene, Michael Maynord, Chinmaya Devaraj, Cornelia Fermuller, Yiannis Aloimonos

We introduce Egocentric Object Manipulation Graphs (Ego-OMG) - a novel representation for activity modeling and anticipation of near future actions integrating three components: 1) semantic temporal structure of activities, 2) short-term dynamics, and 3) representations for appearance. Semantic temporal structure is modeled through a graph, embedded through a Graph Convolutional Network, whose states model characteristics of and relations between hands and objects. These state representations derive from all three levels of abstraction, and span segments delimited by the making and breaking of hand-object contact. Short-term dynamics are modeled in two ways: A) through 3D convolutions, and B) through anticipating the spatiotemporal end points of hand trajectories, where hands come into contact with objects. Appearance is modeled through deep spatiotemporal features produced through existing methods. We note that in Ego-OMG it is simple to swap these appearance features, and thus Ego-OMG is complementary to most existing action anticipation methods. We evaluate Ego-OMG on the EPIC Kitchens Action Anticipation Challenge. The consistency of the egocentric perspective of EPIC Kitchens allows for the utilization of the hand-centric cues upon which Ego-OMG relies. We demonstrate state-of-the-art performance, outranking all other previous published methods by large margins and ranking first on the unseen test set and second on the seen test set of the EPIC Kitchens Action Anticipation Challenge. We attribute the success of Ego-OMG to the modeling of semantic structure captured over long timespans. We evaluate the design choices made through several ablation studies. Code will be released upon acceptance

📄 PDF Abstract BibTeX arXiv:2006.03201

Code (0)

등록된 구현이 없습니다.

Tasks

Action AnticipationAttributeObject

Similar Papers 제목 키워드 기반

EgoAERO: Learning Dexterous Manipulation from a Single Egocentric Video without Object Assets

2026-06-06 · Yichen Niu, Haoran Lv, Xinrui Zhang, Xueyao Wan 외 arxiv

Egocentric RGB-D videos offer a natural source of human dexterous manipulation demonstrations, but existing data is difficult to use for robot learning because object pose, geometry, and contact information are often mis…

Object Tracking

Pandora: Articulated 3D Scene Graphs from Egocentric Vision

2026-03-30 · Alan Yu, Yun Chang, Christopher Xie, Luca Carlone arxiv

Robotic mapping systems typically approach building metric-semantic scene representations from the robot's own sensors and cameras. However, these "first person" maps inherit the robot's own limitations due to its embodi…

Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

2025-01-01 · CVPR 2025 1 · Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori

Learning to use tools or objects in common scenes, particularly handling them in various ways as instructed, is a key challenge for developing interactive robots. Training models to generate such manipulation traject…

valid

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

2026-06-20 · Yangtao Chen, Zixuan Chen, Peiyang Wang, Yong-Lu Li 외 arxiv

Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well aligned with r…

MAPLE: Encoding Dexterous Robotic Manipulation Priors Learned From Egocentric Videos

2025-04-08 · Alexey Gavryushin, Xi Wang, Robert J. S. Malate, Chenyu Yang 외

Large-scale egocentric video datasets capture diverse human activities across a wide range of scenarios, offering rich and detailed insights into how humans interact with objects, especially those that require fine-grain…