paper-with-me

홈 › Papers

SocialMirror: Reconstructing 3D Human Interaction Behaviors from Monocular Videos with Semantic and Geometric Guidance

2026-04-15 · Qi Xia, Peishan Cong, Ziyi Wang, Yujing Sun, Qin Sun, Xinge Zhu, Mao Ye, Ruigang Yang, Yuexin Ma arxiv

Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, precise motion analysis in sports, and natural collaborative behavior in human-robot tasks. Reliable reconstruction in these contexts significantly enhances the realism and effectiveness of AI-driven interactive applications. However, human reconstruction from monocular videos in close-interaction scenarios remains challenging due to severe mutual occlusions, leading local motion ambiguity, disrupted temporal continuity and spatial relationship error. In this paper, we propose SocialMirror, a diffusion-based framework that integrates semantic and geometric cues to effectively address these issues. Specifically, we first leverage high-level interaction descriptions generated by a vision-language model to guide a semantic-guided motion infiller, hallucinating occluded bodies and resolving local pose ambiguities. Next, we propose a sequence-level temporal refiner that enforces smooth, jitter-free motions, while incorporating geometric constraints during sampling to ensure plausible contact and spatial relationships. Evaluations on multiple interaction benchmarks show that SocialMirror achieves state-of-the-art performance in reconstructing interactive human meshes, demonstrating strong generalization across unseen datasets and in-the-wild scenarios. The code will be released upon publication.

📄 PDF Abstract BibTeX arXiv:2604.13581

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos

2026-05-16 · Lixin Xue, Chengwei Zheng, Georgios Paschalidis, Chen Guo 외 arxiv

Reconstructing people, objects, and their interactions in 3D is a long-standing goal for intelligent systems. Often the input is RGB video from a moving camera, making the task ill-posed; depth is ambiguous, humans and o…

Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

2026-05-14 · Yubo Zhao, Yujin Chai, Yunao Dong, Chengfeng Zhao 외 arxiv

Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can reconstruct temporally coherent human an…

Reconstructing Interacting Hands with Interaction Prior from Monocular Images

2023-08-27 · ICCV 2023 1 · Binghui Zuo, Zimeng Zhao, Wenqian Sun, Wei Xie 외

Reconstructing interacting hands from monocular images is indispensable in AR/VR applications. Most existing solutions rely on the accurate localization of each skeleton joint. However, these methods tend to be unreliabl…

4D Human Body Capture from Egocentric Video via 3D Scene Grounding

2020-11-26 · Miao Liu, Dexin Yang, Yan Zhang, Zhaopeng Cui 외

We introduce a novel task of reconstructing a time series of second-person 3D human body meshes from monocular egocentric videos. The unique viewpoint and rapid embodied camera motion of egocentric videos raise additiona…

Time SeriesTime Series Analysis

D$^3$-Human: Dynamic Disentangled Digital Human from Monocular Video

2025-01-03 · Honghu Chen, Bo Peng, Yunfan Tao, Juyong Zhang

We introduce D$^3$-Human, a method for reconstructing Dynamic Disentangled Digital Human geometry from monocular videos. Past monocular video human reconstruction primarily focuses on reconstructing undecoupled clothed h…