paper-with-me

Papers

EgoExo-WM: Unlocking Exo Video for Ego World Models

2026-05-14 · Danny Tran, Roberto Martín-Martín, Kristen Grauman arxiv

Egocentric world models present a promising direction for enabling agents to predict and plan, but their performance is constrained by the limited availability of egocentric training data and its inherent partial observability of humans' physical actions. In contrast, exocentric video is abundant and reveals body poses well, but lacks direct alignment with an agent's action space -- and is not egocentric. We propose a method to bridge this gap by extracting structured body pose from exocentric video as a representation of action and transforming the exocentric video to egocentric video, informed by a human kinematics prior. This process unlocks the integration of in-the-wild exocentric data for egocentric world model training. We show that training whole-body action-conditioned egocentric world models with our converted data significantly improves both prediction quality and downstream planning performance, where we infer the sequence of body poses needed to achieve a visual goal state. Our approach paves the way to enlist arbitrary in-the-wild videos for building powerful egocentric world models, furthering applications in robot planning and augmented-reality guidance.

📄 PDF Abstract BibTeX arXiv:2605.15477

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World

2024-03-24 · CVPR 2024 1 · Yifei HUANG, Guo Chen, Jilan Xu, Mingfang Zhang 외

Being able to map the activities of others into one's own point of view is one fundamental human skill even from a very early age. Taking a step toward understanding this human ability, we introduce EgoExoLearn, a large-…

Action AnticipationAction Quality AssessmentLong Term AnticipationVideo Retrieval

EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding

2024-06-13 · Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng 외

We present EgoExo-Fitness, a new full-body action understanding dataset, featuring fitness sequence videos recorded from synchronized egocentric and fixed exocentric (third-person) cameras. Compared with existing full-bo…

Action ClassificationAction LocalizationAction Understanding

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

2025-10-30 · Minjoon Jung, Junbin Xiao, Junghyun Kim, Byoung-Tak Zhang 외 arxiv

Do Video-LLMs have consistent temporal understanding when videos capture the same event from different viewpoints? To study this question, we introduce EgoExo-Con(sistency), a benchmark of synchronized egocentric and exo…

Reinforcement Learning

EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos

2025-04-16 · Jilan Xu, Yifei HUANG, Baoqi Pei, Junlin Hou 외

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an…

PredictionVideo Prediction

Exo2EgoSyn: Unlocking Foundation Video Generation Models for Exocentric-to-Egocentric Video Synthesis

2025-11-25 · Mohammad Mahdi, Yuqian Fu, Nedko Savov, Jiancheng Pan 외 arxiv

Foundation video generation models such as WAN 2.2 exhibit strong text- and image-conditioned synthesis abilities but remain constrained to the same-view generation setting. In this work, we introduce Exo2EgoSyn, an adap…

Video Generation