paper-with-me

홈 › Papers

High-Fidelity 4D Hand-Object Capture via Multi-View Spatiotemporal Tracking and Physics-Aware Gaussians

2026-06-14 · Bo Peng, Xu Chen, Yi Gu, Hidenobu Matsuki, Mingsong Dou, Jingjing Shen, Deying Kong, Juyong Zhang, Zhengyang Shen arxiv

The growing demand for high-fidelity 4D hand-object interaction (HOI) data in embodied AI and spatial computing is currently bottlenecked by the reliance on pre-scanned object templates and physical markers. While recent methods have demonstrated promising results in reconstructing 4D hand-object interaction from videos, they are highly sensitive to initial estimates of hand and object poses. Yet, estimating these poses from images is challenging, in particular under severe occlusion which is inherent in hand-object interaction scenarios. We propose a novel system for the robust and accurate reconstruction of hands and objects from synchronized and calibrated multi-view videos without requiring any templates or markers. Our system consists of two main components with key innovations: (1) a multi-view feed-forward transformer model that aggregates cross-view geometry and temporal cues to provide a reliable, metric-consistent initialization for both poses and dense object geometry, and (2) a hand-object physics-aware Gaussian-based optimization framework to refine the initial estimates, integrating tetrahedral constraints, collision refinement, and appearance decomposition to produce physically plausible and visually accurate reconstruction. Validated on public benchmarks and an extensive internal dataset, our pipeline achieves highly robust, artifact-free reconstruction, providing an efficient foundation for automated 4D asset generation. Our project page are available at https://zyshen021.github.io/HOSTPG/.

📄 PDF Abstract BibTeX arXiv:2606.15908

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MANUS: Markerless Grasp Capture using Articulated 3D Gaussians

2023-12-04 · CVPR 2024 1 · Chandradeep Pokhariya, Ishaan N Shah, Angela Xing, Zekun Li 외

Understanding how we grasp objects with our hands has important applications in areas like robotics and mixed reality. However, this challenging problem requires accurate modeling of the contact between hands and objects…

Mixed RealityObject

HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping

2026-04-16 · Jongbin Lim, Taeyun Ha, Mingi Choi, Jisoo Kim 외 arxiv

We present HRDexDB, a paired cross-embodiment dexterous grasping dataset of high-fidelity dexterous grasping sequences featuring both human and diverse robotic hands. Unlike existing datasets, HRDexDB provides a comprehe…

VEPHand: View-Efficient Photometric Hand Performance Capture at Scale

2026-06-14 · Zhengyang Shen, Kai-Hung Chang, Erroll Wood, Deying Kong 외 arxiv

Robust, high-fidelity 3D hand capture, while fundamental to digital human creation, remains challenging with practical multi-view systems that balance rich photometry with the geometric ambiguities of reconstruction aris…

DressRecon: Freeform 4D Human Reconstruction from Monocular Video

2024-09-30 · Jeff Tan, Donglai Xiang, Shubham Tulsiani, Deva Ramanan 외

We present a method to reconstruct time-consistent human body models from monocular videos, focusing on extremely loose clothing or handheld object interactions. Prior work in human reconstruction is either limited to ti…

ObjectOptical Flow Estimation

Ins-HOI: Instance Aware Human-Object Interactions Recovery

2023-12-15 · Jiajun Zhang, Yuxiang Zhang, Hongwen Zhang, Xiao Zhou 외

Accurately modeling detailed interactions between human/hand and object is an appealing yet challenging task. Current multi-view capture systems are only capable of reconstructing multiple subjects into a single, unified…

DescriptiveDisentanglementHuman-Object Interaction DetectionObject+1