paper-with-me

홈 › Papers

Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints

2026-03-12 · Chenyangguang Zhang, Botao Ye, Boqi Chen, Alexandros Delitzas, Fangjinhua Wang, Marc Pollefeys, Xi Wang arxiv

Controllable video generation for complex hand-object interactions is a critical step toward building visual world models. However, existing methods often struggle to achieve fine-grained, 3D-consistent hand articulation in generated videos. By relying on dense 2D trajectories or implicit pose representations, they collapse crucial geometric structures into spatially ambiguous signals, leading to severe motion inconsistencies and hallucinated artifacts under egocentric occlusions. To address this, we propose leveraging sparse 3D hand joints as explicit control signals with three key advantages: explicit geometry to resolve occlusions, an intuitive interface for interactive editing, and cross-embodiment generalization to robotic hands. Built upon this, our efficient control module extracts occlusion-aware features from the source reference frame by penalizing unreliable visual features from hidden joints, and employs a 3D-based weighting mechanism to handle dynamically occluded target joints during motion propagation. Meanwhile, it directly injects 3D geometric embeddings into the latent space to enforce structural consistency. To facilitate robust training and evaluation, we develop an automated annotation pipeline, yielding 1M high-quality egocentric video clips paired with precise hand trajectories. Experiments demonstrate that our approach outperforms state-of-the-art baselines, generating high-fidelity egocentric videos with realistic hand-object interactions.

📄 PDF Abstract BibTeX arXiv:2603.11755

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control

2026-05-25 · Qiao Gu, Lingni Ma, Adam W Harley, Richard Newcombe 외 arxiv

Controllable and physically grounded egocentric video generation is essential for embodied agents to reason about how their own and others' actions manifest and change the world. Compared to generic video synthesis, egoc…

Video Generation

EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses

2025-11-22 · Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy 외 arxiv

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controll…

Video GenerationVideo Prediction

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

2026-07-30 · Zexuan Yan, Yuzhou Wu, Yue Ma, Zonghang He 외 arxiv

Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action…

Video Generation

EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation

2026-04-01 · Abhishek Saroha, Huajian Zeng, Xingxing Zuo, Daniel Cremers 외 arxiv

Understanding and predicting object motion from egocentric video is fundamental to embodied perception and interaction. However, generating physically consistent 6DoF trajectories remains challenging due to occlusions, f…

Collision Avoidance

Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos

2024-06-13 · Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue 외

Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the v…

Audio GenerationRetrieval-augmented Generation