paper-with-me

Papers

UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation

2025-08-02 · Chaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka, Juan Carlos Niebles, Ehsan Adeli arxiv

Egocentric human motion generation and forecasting with scene-context is crucial for enhancing AR/VR experiences, improving human-robot interaction, advancing assistive technologies, and enabling adaptive healthcare solutions by accurately predicting and simulating movement from a first-person perspective. However, existing methods primarily focus on third-person motion synthesis with structured 3D scene contexts, limiting their effectiveness in real-world egocentric settings where limited field of view, frequent occlusions, and dynamic cameras hinder scene perception. To bridge this gap, we introduce Egocentric Motion Generation and Egocentric Motion Forecasting, two novel tasks that utilize first-person images for scene-aware motion synthesis without relying on explicit 3D scene. We propose UniEgoMotion, a unified conditional motion diffusion model with a novel head-centric motion representation tailored for egocentric devices. UniEgoMotion's simple yet effective design supports egocentric motion reconstruction, forecasting, and generation from first-person visual inputs in a unified framework. Unlike previous works that overlook scene semantics, our model effectively extracts image-based scene context to infer plausible 3D motion. To facilitate training, we introduce EE4D-Motion, a large-scale dataset derived from EgoExo4D, augmented with pseudo-ground-truth 3D motion annotations. UniEgoMotion achieves state-of-the-art performance in egocentric motion reconstruction and is the first to generate motion from a single egocentric image. Extensive evaluations demonstrate the effectiveness of our unified framework, setting a new benchmark for egocentric motion modeling and unlocking new possibilities for egocentric applications.

📄 PDF Abstract BibTeX arXiv:2508.01126

Code (0)

등록된 구현이 없습니다.

Tasks

Motion ForecastingMotion Synthesis

Similar Papers 제목 키워드 기반

EgoPriMo: Egocentric Motion Generation for Interactive Humanoid Control

2026-06-07 · Haoyang Ge, Peng Ren, Yukun Shi, Cong Huang 외 arxiv

Humanoid robots require whole-body motions that adapt to scene context, task requirements, and user intent. Motion tracking reproduces specified trajectories, and humanoid vision-language-action systems provide semantic …

Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting

2026-07-17 · Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Xiang Li 외 arxiv

Perceiving multimodal cues and forecasting fine-grained actions from an egocentric (Ego) perspective is vital for applications like robot manipulation. However, previous studies either rely mainly on under-informed visua…

Robot Manipulation

EggHand: A Multimodal Foundation Model for Egocentric Hand Pose Forecasting

2026-05-08 · Jaeyoung Choi, Hyeondong Kim, Yujin Kim, Daehee Park arxiv

Forecasting future 3D hand pose sequences from egocentric video is essential for understanding human intention and enabling embodied applications such as AR/VR assistance and human-robot interaction. However, this task r…

HOIMotion: Forecasting Human Motion During Human-Object Interactions Using Egocentric 3D Object Bounding Boxes

2024-07-02 · Zhiming Hu, Zheming Yin, Daniel Haeufle, Syn Schmitt 외

We present HOIMotion - a novel approach for human motion forecasting during human-object interactions that integrates information about past body poses and egocentric 3D object bounding boxes. Human motion forecasting is…

DecoderHuman-Object Interaction DetectionMotion ForecastingObject

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video

2026-07-20 · Xiaozhong Lyu, Gen Li, Zhiyin Qian, Xucong Zhang 외 hf

Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal…

Depth Estimation