paper-with-me

Papers

StackFLOW: Monocular Human-Object Reconstruction by Stacked Normalizing Flow with Offset

2024-07-30 · Chaofan Huo, Ye Shi, Yuexin Ma, Lan Xu, Jingyi Yu, Jingya Wang

Modeling and capturing the 3D spatial arrangement of the human and the object is the key to perceiving 3D human-object interaction from monocular images. In this work, we propose to use the Human-Object Offset between anchors which are densely sampled from the surface of human mesh and object mesh to represent human-object spatial relation. Compared with previous works which use contact map or implicit distance filed to encode 3D human-object spatial relations, our method is a simple and efficient way to encode the highly detailed spatial correlation between the human and object. Based on this representation, we propose Stacked Normalizing Flow (StackFLOW) to infer the posterior distribution of human-object spatial relations from the image. During the optimization stage, we finetune the human body pose and object 6D pose by maximizing the likelihood of samples based on this posterior distribution and minimizing the 2D-3D corresponding reprojection loss. Extensive experimental results show that our method achieves impressive results on two challenging benchmarks, BEHAVE and InterCap datasets.

📄 PDF Abstract BibTeX arXiv:2407.20545

Code (1)

huochf/StackFLOW 공식 구현 pytorch

Tasks

Human-Object Interaction DetectionObjectObject Reconstruction

Similar Papers 제목 키워드 기반

Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

2026-05-14 · Yubo Zhao, Yujin Chai, Yunao Dong, Chengfeng Zhao 외 arxiv

Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can reconstruct temporally coherent human an…

CoGS: Compositional Dynamic Human-Object Scenes Gaussian Splatting from Monocular Video

2026-06-27 · Jerrin Bright, John Zelek arxiv

Reconstructing dynamic human--object interaction scenes from monocular video is difficult because the human, manipulated object, and background obey different motion models while sharing the same pixels. Existing dynamic…

RobustFusion: Robust Volumetric Performance Reconstruction under Human-object Interactions from Monocular RGBD Stream

2021-04-30 · Zhuo Su, Lan Xu, Dawei Zhong, Zhong Li 외

High-quality 4D reconstruction of human performance with complex interactions to various objects is essential in real-world scenarios, which enables numerous immersive VR/AR applications. However, recent advances still f…

4D reconstructionDisentanglementHuman-Object Interaction DetectionHuman Parsing+3

ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors

2026-03-04 · Zihao Huang, Tianqi Liu, Zhaoxi Chen, Shaocong Xu 외 arxiv

Synthesizing physically plausible articulated human-object interactions (HOI) without 3D/4D supervision remains a fundamental challenge. While recent zero-shot approaches leverage video diffusion models to synthesize hum…

Inverse Rendering

Gravity-Aware Monocular 3D Human-Object Reconstruction

2021-08-19 · ICCV 2021 10 · Rishabh Dabral, Soshi Shimada, Arjun Jain, Christian Theobalt 외

This paper proposes GraviCap, i.e., a new approach for joint markerless 3D human motion capture and object trajectory estimation from monocular RGB videos. We focus on scenes with objects partially observed during a free…

Human-Object Interaction DetectionObjectObject Reconstruction