paper-with-me

홈 › Papers

Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model

2025-03-21 · CVPR 2025 1 · Yingying Fan, Quanwei Yang, Kaisiyuan Wang, Hang Zhou, YingYing Li, Haocheng Feng, Errui Ding, Yu Wu, Jingdong Wang

Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g., objects) have not been well investigated. Despite human hand synthesis already being an intricate problem, generating objects in contact with hands and their interactions presents an even more challenging task, especially when the objects exhibit obvious variations in size and shape. To cope with these issues, we present a novel video Reenactment framework focusing on Human-Object Interaction (HOI) via an adaptive Layout-instructed Diffusion model (Re-HOLD). Our key insight is to employ specialized layout representation for hands and objects, respectively. Such representations enable effective disentanglement of hand modeling and object adaptation to diverse motion sequences. To further improve the generation quality of HOI, we have designed an interactive textural enhancement module for both hands and objects by introducing two independent memory banks. We also propose a layout-adjusting strategy for the cross-object reenactment scenario to adaptively adjust unreasonable layouts caused by diverse object sizes during inference. Comprehensive qualitative and quantitative evaluations demonstrate that our proposed framework significantly outperforms existing methods. Project page: https://fyycs.github.io/Re-HOLD.

📄 PDF Abstract BibTeX arXiv:2503.16942

Code (0)

등록된 구현이 없습니다.

Tasks

DisentanglementHuman-Object Interaction DetectionObjectVideo Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

2025-06-15 · Zhelun Shen, Chenming Wu, Junsheng Zhou, Chen Zhao 외

Digital human video generation is gaining traction in fields like education and e-commerce, driven by advancements in head-body animation and lip-syncing technologies. However, realistic Hand-Object Interaction (HOI) - t…

ObjectVideo Generation

HOReeNet: 3D-aware Hand-Object Grasping Reenactment

2022-11-11 · Changhwa Lee, Junuk Cha, Hansol Lee, Seongyeong Lee 외

We present HOReeNet, which tackles the novel task of manipulating images involving hands, objects, and their interactions. Especially, we are interested in transferring objects of source images to target images and manip…

3D ReconstructionObject

MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model

2026-03-16 · Jinguang Tong, Jinbo Wu, Kaisiyuan Wang, Zhelun Shen 외 arxiv

Human-Object Interaction (HOI) video reenactment aims to transfer the interaction dynamics of a source video to a novel target object while preserving realistic hand-object coordination. Existing methods typically rely o…

Video Generation

GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection

2026-03-06 · Xuan Huang, Mochu Xiang, Zhelun Shen, Jinbo Wu 외 arxiv

Hand-Object Interaction (HOI) remains a core challenge in digital human video synthesis, where models must generate physically plausible contact and preserve object identity across frames. Although recent HOI reenactment…

Video Generation

SelfieAvatar: Real-time Head Avatar reenactment from a Selfie Video

2026-01-26 · Wei Liang, Hui Yu, Derui Ding, Rachael E. Jack 외 arxiv

Head avatar reenactment focuses on creating animatable personal avatars from monocular videos, serving as a foundational element for applications like social signal understanding, gaming, human-machine interaction, and c…

Image Generation