paper-with-me

Papers

EgoGen: An Egocentric Synthetic Data Generator

2024-01-16 · CVPR 2024 1 · Gen Li, Kaifeng Zhao, Siwei Zhang, Xiaozhong Lyu, Mihai Dusmanu, Yan Zhang, Marc Pollefeys, Siyu Tang

Understanding the world in first-person view is fundamental in Augmented Reality (AR). This immersive perspective brings dramatic visual changes and unique challenges compared to third-person views. Synthetic data has empowered third-person-view vision models, but its application to embodied egocentric perception tasks remains largely unexplored. A critical challenge lies in simulating natural human movements and behaviors that effectively steer the embodied cameras to capture a faithful egocentric representation of the 3D world. To address this challenge, we introduce EgoGen, a new synthetic data generator that can produce accurate and rich ground-truth training data for egocentric perception tasks. At the heart of EgoGen is a novel human motion synthesis model that directly leverages egocentric visual inputs of a virtual human to sense the 3D environment. Combined with collision-avoiding motion primitives and a two-stage reinforcement learning approach, our motion synthesis model offers a closed-loop solution where the embodied perception and movement of the virtual human are seamlessly coupled. Compared to previous works, our model eliminates the need for a pre-defined global path, and is directly applicable to dynamic environments. Combined with our easy-to-use and scalable data generation pipeline, we demonstrate EgoGen's efficacy in three tasks: mapping and localization for head-mounted cameras, egocentric camera tracking, and human mesh recovery from egocentric views. EgoGen will be fully open-sourced, offering a practical solution for creating realistic egocentric training data and aiming to serve as a useful tool for egocentric computer vision research. Refer to our project page: https://ego-gen.github.io/.

📄 PDF Abstract BibTeX arXiv:2401.08739

Code (0)

등록된 구현이 없습니다.

Tasks

Human Mesh RecoveryMotion Synthesis

Similar Papers 제목 키워드 기반

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

2026-07-30 · Zexuan Yan, Yuzhou Wu, Yue Ma, Zonghang He 외 arxiv

Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action…

Video Generation

SEED4D: A Synthetic Ego--Exo Dynamic 4D Data Generator, Driving Dataset and Benchmark

2024-12-01 · Marius Kästingschäfer, Théo Gieruc, Sebastian Bernhard, Dylan Campbell 외

Models for egocentric 3D and 4D reconstruction, including few-shot interpolation and extrapolation settings, can benefit from having images from exocentric viewpoints as supervision signals. No existing dataset provides …

2k4D reconstructionAutonomous Driving

Deep Future Gaze: Gaze Anticipation on Egocentric Videos Using Adversarial Networks

2017-07-01 · CVPR 2017 7 · Mengmi Zhang, Keng Teck Ma, Joo Hwee Lim, Qi Zhao 외

We introduce a new problem of gaze anticipation on egocentric videos. This substantially extends the conventional gaze prediction problem to future frames by no longer confining it on the current frame. To solve this pro…

Gaze Prediction

Egocentric Videoconferencing

2021-07-07 · Mohamed Elgharib, Mohit Mendiratta, Justus Thies, Matthias Nießner 외

We introduce a method for egocentric videoconferencing that enables hands-free video calls, for instance by people wearing smart glasses or other mixed-reality devices. Videoconferencing portrays valuable non-verbal comm…

Face ReenactmentMixed Reality

EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation

2026-05-18 · Rosario Leonardi, Francesco Ragusa, Daniele Materia, Alessandro Passanisi 외 arxiv

Collecting large-scale egocentric video datasets with dense spatial and temporal annotations is costly, slow, and often constrained by environmental biases, privacy constraints, and limited coverage of interaction patter…

Active Object DetectionAction SegmentationVideo Generation