paper-with-me

홈 › Papers

HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction

2025-08-22 · Sara Rojas, Matthieu Armando, Bernard Ghamen, Philippe Weinzaepfel, Vincent Leroy, Gregory Rogez arxiv

Recovering the 3D geometry of a scene from a sparse set of uncalibrated images is a long-standing problem in computer vision. While recent learning-based approaches such as DUSt3R and MASt3R have demonstrated impressive results by directly predicting dense scene geometry, they are primarily trained on outdoor scenes with static environments and struggle to handle human-centric scenarios. In this work, we introduce HAMSt3R, an extension of MASt3R for joint human and scene 3D reconstruction from sparse, uncalibrated multi-view images. First, we exploit DUNE, a strong image encoder obtained by distilling, among others, the encoders from MASt3R and from a state-of-the-art Human Mesh Recovery (HMR) model, multi-HMR, for a better understanding of scene geometry and human bodies. Our method then incorporates additional network heads to segment people, estimate dense correspondences via DensePose, and predict depth in human-centric environments, enabling a more comprehensive 3D reconstruction. By leveraging the outputs of our different heads, HAMSt3R produces a dense point map enriched with human semantic information in 3D. Unlike existing methods that rely on complex optimization pipelines, our approach is fully feed-forward and efficient, making it suitable for real-world applications. We evaluate our model on EgoHumans and EgoExo4D, two challenging benchmarks con taining diverse human-centric scenarios. Additionally, we validate its generalization to traditional multi-view stereo and multi-view pose regression tasks. Our results demonstrate that our method can reconstruct humans effectively while preserving strong performance in general 3D reconstruction tasks, bridging the gap between human and scene understanding in 3D vision.

📄 PDF Abstract BibTeX arXiv:2508.16433

Code (0)

등록된 구현이 없습니다.

Tasks

Human Mesh RecoveryScene Understanding3D Reconstruction

Similar Papers 제목 키워드 기반

EATR-Stereo: Embodiment-Aware Token Routing of Paired Stereo Evidence for Humanoid Vision-Language-Action Control

2026-08-18 · Songwei Wu, Rui Zhao, Fan Yang, Zhongqiang Nie 외 arxiv

Long-horizon humanoid vision--language--action (VLA) control with head-mounted stereo cameras requires visual interfaces that can exploit complementary views while maintaining compatibility with pretrained representation…

StereoGS: Sparse-View 3D Gaussian Splatting via Stereo Priors

2026-06-29 · Wenhao Yuan, Yiyuan Ge, Deli Cai arxiv

3D Gaussian Splatting (3DGS) has achieved remarkable success in real-time novel view synthesis, yet it suffers from severe overfitting under sparse-view settings due to insufficient geometric constraints. While recent me…

Novel View SynthesisDepth Estimation

DiffuStereo: High Quality Human Reconstruction via Diffusion-based Stereo Using Sparse Cameras

2022-07-16 · Ruizhi Shao, Zerong Zheng, Hongwen Zhang, Jingxiang Sun 외

We propose DiffuStereo, a novel system using only sparse cameras (8 in this work) for high-quality 3D human reconstruction. At its core is a novel diffusion-based stereo module, which introduces diffusion models, a type …

3D Human Reconstruction4kDepth EstimationStereo Matching

Splat-SAP: Feed-Forward Gaussian Splatting for Human-Centered Scene with Scale-Aware Point Map Reconstruction

2025-11-27 · Boyao Zhou, Shunyuan Zheng, Zhanfeng Liao, Zihan Ma 외 arxiv

We present Splat-SAP, a feed-forward approach to render novel views of human-centered scenes from binocular cameras with large sparsity. Gaussian Splatting has shown its promising potential in rendering tasks, but it typ…

PSMNet: Position-aware Stereo Merging Network for Room Layout Estimation

2022-03-30 · CVPR 2022 1 · HaiYan Wang, Will Hutchcroft, Yuguang Li, Zhiqiang Wan 외

In this paper, we propose a new deep learning-based method for estimating room layout given a pair of 360 panoramas. Our system, called Position-aware Stereo Merging Network or PSMNet, is an end-to-end joint layout-pose …

Omnnidirectional Stereo Depth EstimationPositionRoom Layout Estimation