paper-with-me

홈 › Papers

G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis

2024-04-18 · CVPR 2024 1 · Yufei Ye, Abhinav Gupta, Kris Kitani, Shubham Tulsiani

We propose G-HOP, a denoising diffusion based generative prior for hand-object interactions that allows modeling both the 3D object and a human hand, conditioned on the object category. To learn a 3D spatial diffusion model that can capture this joint distribution, we represent the human hand via a skeletal distance field to obtain a representation aligned with the (latent) signed distance field for the object. We show that this hand-object prior can then serve as generic guidance to facilitate other tasks like reconstruction from interaction clip and human grasp synthesis. We believe that our model, trained by aggregating seven diverse real-world interaction datasets spanning across 155 categories, represents a first approach that allows jointly generating both hand and object. Our empirical evaluations demonstrate the benefit of this joint prior in video-based reconstruction and human grasp synthesis, outperforming current task-specific baselines. Project website: https://judyye.github.io/ghop-www

📄 PDF Abstract BibTeX arXiv:2404.12383

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingObject

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos

2026-02-25 · Yufei Ye, Jiaman Li, Ryan Rong, C. Karen Liu arxiv

Egocentric manipulation videos are highly challenging due to severe occlusions during interactions and frequent object entries and exits from the camera view as the person moves. Current methods typically focus on recove…

Pose Estimation

CHOIR: Contact-aware 4D Hand-Object Interaction Reconstruction

2026-05-20 · Hao Xu, Yilin Liu, Yinqiao Wang, Chi-Wing Fu 외 arxiv

We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape with 6D pose over time, and the when/where of contact. Such a capability wo…

BG-HOP: A Bimanual Generative Hand-Object Prior

2025-06-08 · Sriram Krishna, Sravan Chittupalli, Sungjae Park

In this work, we present BG-HOP, a generative prior that seeks to model bimanual hand-object interactions in 3D. We address the challenge of limited bimanual interaction data by extending existing single-hand generative …

MagicHOI: Leveraging 3D Priors for Accurate Hand-object Reconstruction from Short Monocular Video Clips

2025-08-07 · Shibo Wang, Haonan He, Maria Parelli, Christoph Gebhardt 외 arxiv

Most RGB-based hand-object reconstruction methods rely on object templates, while template-free methods typically assume full object visibility. This assumption often breaks in real-world settings, where fixed camera vie…

Novel View Synthesis

Affordance-Guided Diffusion Prior for 3D Hand Reconstruction

2025-10-01 · Naru Suzuki, Takehiko Ohkawa, Tatsuro Banno, Jihyun Lee 외 arxiv

How can we reconstruct 3D hand poses when large portions of the hand are heavily occluded by itself or by objects? Humans often resolve such ambiguities by leveraging contextual knowledge -- such as affordances, where an…

Hand Pose Estimation