paper-with-me

홈 › Papers

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

2026-04-21 · Xiangyang Luo, Xiaozhe Xin, Tao Feng, Xu Guo, Meiguang Jin, Junfeng Ma arxiv

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capability, still frequently fail on (i) the structural stability of sensitive regions such as hands and faces and (ii) physically plausible contact (e.g., avoiding hand--object interpenetration). We present CoInteract, an end-to-end framework for HOI video synthesis conditioned on a person reference image, a product reference image, text prompts, and speech audio. CoInteract introduces two complementary designs embedded into a Diffusion Transformer (DiT) backbone. First, we propose a Human-Aware Mixture-of-Experts (MoE) that routes tokens to lightweight, region-specialized experts via spatially supervised routing, improving fine-grained structural fidelity with minimal parameter overhead. Second, we propose Spatially-Structured Co-Generation, a dual-stream training paradigm that jointly models an RGB appearance stream and an auxiliary HOI structure stream to inject interaction geometry priors. During training, the HOI stream attends to RGB tokens and its supervision regularizes shared backbone weights; at inference, the HOI branch is removed for zero-overhead RGB generation. Experimental results demonstrate that CoInteract significantly outperforms existing methods in structural stability, logical consistency, and interaction realism.

📄 PDF Abstract BibTeX arXiv:2604.19636

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions

2026-05-28 · Omer Benishu, Gal Fiebelman, Sagie Benaim arxiv

We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target object represented as 3D Gaussian Splats (3DGS), our goal is to synthesize…

Weave: Learning Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions

2026-09-15 · Liu Cao, Xingze Wu, Jingzhi Cui, Botian Xu 외 arxiv

Learning humanoid-object interaction requires coordinating whole-body balance, locomotion, and dexterous hand contact to control both robot and object motion. Human demonstrations provide examples of coordinated interact…

Recovering Physically Plausible Human-Object Interactions from Monocular Videos

2026-06-03 · Dingbang Huang, Etienne Vouga, Qixing Huang, Georgios Pavlakos arxiv

In this paper, we propose RePHO, a method to reconstruct physically plausible human-object interactions (HOI) from monocular videos. While existing kinematic-based approaches produce visually plausible motion, they often…

Reinforcement Learning

ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors

2026-03-04 · Zihao Huang, Tianqi Liu, Zhaoxi Chen, Shaocong Xu 외 arxiv

Synthesizing physically plausible articulated human-object interactions (HOI) without 3D/4D supervision remains a fundamental challenge. While recent zero-shot approaches leverage video diffusion models to synthesize hum…

Inverse Rendering

PIAvatar: Physically Interactive Avatars via Deformation Gradient Decoupling

2026-06-19 · Sang-Hun Han, Min-Gyu Park, Jisu Shin, Seunghyun Shin 외 arxiv

3D human avatars have shown impressive visual fidelity driven by pose-conditioned models, yet they still lack the physical ability required for interactions with each other and environments. Although recent studies have …