paper-with-me

Papers

PhotoHOI: Synthesizing 3D Hand-Object Interactions from a Single RGB Photograph

2026-08-03 · Zhenhao Zhang, Jiajun Zhang, Wei Min, Yebin Liu arxiv

Hand-object interaction (HOI) is a fundamental human behavior with broad applications in AR/VR, digital humans, and embodied interaction. Existing methods typically require predefined object geometry, object trajectories, or task-specific conditions, limiting their use with natural real-world inputs. To address this, we study a more practical problem of synthesizing 3D hand-object interaction sequences from a single RGB photograph and an open-vocabulary language instruction, and introduce PhotoHOI. PhotoHOI first uses a vision-language model to parse the input image and instruction into a structured task specification, including the interaction object, target region, and spatial relation. It then recovers a compact task-relevant 3D scene and plans a smooth collision-aware object trajectory based on the recovered object states, support relations, and surrounding scene geometry. To synthesize hand motion that generalizes to real-world photographs and unseen objects, it learns transferable task-conditioned contact and contact-conditioned grasp priors from large-scale affordance and HOI data. The grasp is further refined in a learned latent space, constraining the optimization to a plausible hand-pose manifold. Experiments on GRAB and H2O demonstrate improved contact quality and reduced penetration over representative baselines. Results on real-world photographs further demonstrate higher task success and scene consistency, together with generalization to unseen objects and open-vocabulary instructions.

📄 PDF Abstract BibTeX arXiv:2608.01905

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Affordance Diffusion: Synthesizing Hand-Object Interactions

2023-03-21 · CVPR 2023 1 · Yufei Ye, Xueting Li, Abhinav Gupta, Shalini De Mello 외

Recent successes in image synthesis are powered by large-scale diffusion models. However, most methods are currently limited to either text- or image-conditioned generation for synthesizing an entire image, texture trans…

DescriptiveImage GenerationObject

DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions

2025-09-26 · Quanzhou Li, Zhonghua Wu, Jingbo Wang, Chen Change Loy 외 arxiv

Learning to generate dual-hand grasps that respect object semantics is essential for robust hand-object interaction but remains largely underexplored due to dataset scarcity. Existing grasp datasets predominantly focus o…

How Do I Do That? Synthesizing 3D Hand Motion and Contacts for Everyday Interactions

2025-04-16 · CVPR 2025 1 · Aditya Prakash, Benjamin Lundell, Dmitry Andreychuk, David Forsyth 외

We tackle the novel problem of predicting 3D hand motion and contact maps (or Interaction Trajectories) given a single RGB view, action text, and a 3D contact point on the object as input. Our approach consists of (1) In…

DecoderDiversity

COUCH: Towards Controllable Human-Chair Interactions

2022-05-01 · Xiaohan Zhang, Bharat Lal Bhatnagar, Vladimir Guzov, Sebastian Starke 외

Humans interact with an object in many different ways by making contact at different locations, creating a highly complex motion space that can be difficult to learn, particularly when synthesizing such human interaction…

Human-Object Interaction DetectionObject

GraspDiffusion: Synthesizing Realistic Whole-body Hand-Object Interaction

2024-10-17 · Patrick Kwon, Hanbyul Joo

Recent generative models can synthesize high-quality images but often fail to generate humans interacting with objects using their hands. This arises mostly from the model's misunderstanding of such interactions, and the…

Human-Object Interaction DetectionImage GenerationObject