paper-with-me

홈 › Papers

Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos

2024-03-11 · Mi Luo, Zihui Xue, Alex Dimakis, Kristen Grauman

We investigate exocentric-to-egocentric cross-view translation, which aims to generate a first-person (egocentric) view of an actor based on a video recording that captures the actor from a third-person (exocentric) perspective. To this end, we propose a generative framework called Exo2Ego that decouples the translation process into two stages: high-level structure transformation, which explicitly encourages cross-view correspondence between exocentric and egocentric views, and a diffusion-based pixel-level hallucination, which incorporates a hand layout prior to enhance the fidelity of the generated egocentric view. To pave the way for future advancements in this field, we curate a comprehensive exo-to-ego cross-view translation benchmark. It consists of a diverse collection of synchronized ego-exo tabletop activity video pairs sourced from three public datasets: H2O, Aria Pilot, and Assembly101. The experimental results validate that Exo2Ego delivers photorealistic video results with clear hand manipulation details and outperforms several baselines in terms of both synthesis quality and generalization ability to new actions.

📄 PDF Abstract BibTeX arXiv:2403.06351

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationTranslation

Methods 이 논문이 사용한 방법론

ARiA 설명 없음

Similar Papers 제목 키워드 기반

Attention-Propagation Network for Egocentric Heatmap to 3D Pose Lifting

2024-02-28 · CVPR 2024 1 · Taeho Kang, Youngki Lee

We present EgoTAP, a heatmap-to-3D pose lifting method for highly accurate stereo egocentric 3D pose estimation. Severe self-occlusion and out-of-view limbs in egocentric camera views make accurate pose estimation a chal…

3D Pose EstimationEgocentric Pose EstimationPose EstimationPosition

EFM3D: A Benchmark for Measuring Progress Towards 3D Egocentric Foundation Models

2024-06-14 · Julian Straub, Daniel DeTone, Tianwei Shen, Nan Yang 외

The advent of wearable computers enables a new source of context for AI that is embedded in egocentric sensor data. This new egocentric data comes equipped with fine-grained 3D location information and thus presents the …

3D Object Detection3D ReconstructionMulti-View 3D Reconstructionobject-detection+1

Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations

2025-03-25 · CVPR 2025 1 · Jungin Park, Jiyoung Lee, Kwanghoon Sohn

View-invariant representation learning from egocentric (first-person, ego) and exocentric (third-person, exo) videos is a promising approach toward generalizing video understanding systems across multiple viewpoints. How…

Representation LearningVideo Understanding

EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes

2026-07-01 · Jihyeok Jung, Jeewu Lee, Sanghyeop Kim, Chanhee Han 외 arxiv

Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to evaluate egocentric perspective itself in isolation. However, understanding first-…

Scene Understanding

WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation

2025-11-27 · Quanjian Song, Yiren Song, Kelly Peng, Yuan Gao 외 arxiv

Recent advances in video world models enable interactive environments with free navigation, making translation between first-person (egocentric) and third-person (exocentric) perspectives increasingly important. However,…

Video Generation