paper-with-me

홈 › Papers

SHARE: Scene-Human Aligned Reconstruction

2025-10-17 · Joshua Li, Brendan Chharawala, Chang Shu, Xue Bin Peng, Pengcheng Xi arxiv

Animating realistic character interactions with the surrounding environment is important for autonomous agents in gaming, AR/VR, and robotics. However, current methods for human motion reconstruction struggle with accurately placing humans in 3D space. We introduce Scene-Human Aligned REconstruction (SHARE), a technique that leverages the scene geometry's inherent spatial cues to accurately ground human motion reconstruction. Each reconstruction relies solely on a monocular RGB video from a stationary camera. SHARE first estimates a human mesh and segmentation mask for every frame, alongside a scene point map at keyframes. It iteratively refines the human's positions at these keyframes by comparing the human mesh against the human point map extracted from the scene using the mask. Crucially, we also ensure that non-keyframe human meshes remain consistent by preserving their relative root joint positions to keyframe root joints during optimization. Our approach enables more accurate 3D human placement while reconstructing the surrounding scene, facilitating use cases on both curated datasets and in-the-wild web videos. Extensive experiments demonstrate that SHARE outperforms existing methods.

📄 PDF Abstract BibTeX arXiv:2510.15342

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos

2026-06-01 · Jinpeng Liu, Yukang Xu, Yutong Li, Xingyu Liu arxiv

Reconstructing humans and their surrounding environments in a globally consistent 4D space is essential for comprehensive perception. However, prior works typically assume single-view inputs or decouple humans, scenes, a…

Spatial Reasoning

UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video

2026-04-21 · Tanuj Sur, Shashank Tripathi, Nikos Athanasiou, Ha Linh Nguyen 외 arxiv

We introduce UniCon3R, a unified feed-forward framework for online human-scene 4D reconstruction from monocular video. Current feed-forward human-scene reconstruction methods suffer from artifacts, where bodies float abo…

Scene and Human in One World: Reconstruction in a Feedforward Pass

2026-06-26 · Boao Shi, Qiao Feng, Yiming Huang, Lingjie Liu arxiv

Reconstructing humans in dynamic scenes from moving monocular cameras remains challenging due to scale ambiguity, human-scene misalignment, and occlusion interference. Rather than treating human mesh recovery and scene r…

Human Mesh Recovery

AniPixel: Towards Animatable Pixel-Aligned Human Avatar

2023-02-07 · Jinlong Fan, Jing Zhang, Zhi Hou, DaCheng Tao

Although human reconstruction typically results in human-specific avatars, recent 3D scene reconstruction techniques utilizing pixel-aligned features show promise in generalizing to new scenes. Applying these techniques …

3D Scene Reconstruction

G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing

2026-05-26 · Bharath Raj Nagoor Kani, Noah Snavely arxiv

Modern feed-forward 3D reconstruction methods like VGGT predict pixel-aligned pointmaps in camera-centric coordinate frames. However, this choice of coordinate frame is not always optimal. We propose instead to predict p…

3D Reconstruction