paper-with-me

홈 › Papers

OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction

2025-12-22 · Kyungwon Cho, Hanbyul Joo arxiv

The proliferation of commercial egocentric devices offers a unique lens into human behavior, yet reconstructing full-body 3D motion remains difficult due to frequent self-occlusion and the 'out-of-sight' nature of the wearer's limbs. While head and hand trajectories provide sparse anchor points, current methods often overfit to specific hardware optics or rely on expensive, post-hoc optimizations that compromise motion naturalness. In this paper, we present OmniEgoCap, a unified diffusion framework that scales egocentric reconstruction to diverse capture setups. By shifting from short-term windowed estimation to sequence-level inference, our method captures a global perspective and recovers invariant physical attributes, such as height and body proportions, that provide critical constraints for disambiguating head-only cues. To ensure hardware-agnostic generalization, we introduce a geometry-aware visibility augmentation strategy that treats intermittent hand appearances as principled geometric constraints rather than missing data. Our architecture jointly predicts temporally coherent motion and consistent body shape, establishing a new state-of-the-art on public benchmarks and demonstrating robust performance across diverse, in-the-wild environments.

📄 PDF Abstract BibTeX arXiv:2512.19283

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception

2023-06-10 · ICCV 2023 1 · Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters 외

We introduce the Aria Digital Twin (ADT) - an egocentric dataset captured using Aria glasses with extensive object, environment, and human level ground truth. This ADT release contains 200 sequences of real-world activit…

3D Object DetectionBenchmarkingObjectobject-detection+2

Delving Into Egocentric Actions

2015-06-01 · CVPR 2015 6 · Yin Li, Zhefan Ye, James M. Rehg

We address the challenging problem of recognizing the camera wearer's actions from videos captured by an egocentric camera. Egocentric videos encode a rich set of signals regarding the camera wearer, including head movem…

Action RecognitionTemporal Action Localization

HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos

2025-01-06 · CVPR 2025 1 · Jinglei Zhang, Jiankang Deng, Chao Ma, Rolandos Alexandros Potamias

Despite the advent in 3D hand pose estimation, current methods predominantly focus on single-image 3D hand reconstruction in the camera frame, overlooking the world-space motion of the hands. Such limitation prohibits th…

3D Hand Pose EstimationHand Pose EstimationPose Estimation

Egocentric Video Description based on Temporally-Linked Sequences

2017-04-07 · Marc Bolaños, Álvaro Peris, Francisco Casacuberta, Sergi Soler 외

Egocentric vision consists in acquiring images along the day from a first person point-of-view using wearable cameras. The automatic analysis of this information allows to discover daily patterns for improving the qualit…

DecoderVideo Description

Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos

2024-11-14 · Chengbo Yuan, Geng Chen, Li Yi, Yang Gao

Egocentric videos provide valuable insights into human interactions with the physical world, which has sparked growing interest in the computer vision and robotics communities. A critical challenge in fully understanding…

4D reconstructionSelf-Supervised LearningZero-shot Generalization