paper-with-me

홈 › Papers

Reconstructing People, Places, and Cameras

2024-12-23 · CVPR 2025 1 · Lea Müller, Hongsuk Choi, Anthony Zhang, Brent Yi, Jitendra Malik, Angjoo Kanazawa

We present "Humans and Structure from Motion" (HSfM), a method for jointly reconstructing multiple human meshes, scene point clouds, and camera parameters in a metric world coordinate system from a sparse set of uncalibrated multi-view images featuring people. Our approach combines data-driven scene reconstruction with the traditional Structure-from-Motion (SfM) framework to achieve more accurate scene reconstruction and camera estimation, while simultaneously recovering human meshes. In contrast to existing scene reconstruction and SfM methods that lack metric scale information, our method estimates approximate metric scale by leveraging a human statistical model. Furthermore, it reconstructs multiple human meshes within the same world coordinate system alongside the scene point cloud, effectively capturing spatial relationships among individuals and their positions in the environment. We initialize the reconstruction of humans, scenes, and cameras using robust foundational models and jointly optimize these elements. This joint optimization synergistically improves the accuracy of each component. We compare our method to existing approaches on two challenging benchmarks, EgoHumans and EgoExo4D, demonstrating significant improvements in human localization accuracy within the world coordinate frame (reducing error from 3.51m to 1.04m in EgoHumans and from 2.9m to 0.56m in EgoExo4D). Notably, our results show that incorporating human data into the SfM pipeline improves camera pose estimation (e.g., increasing RRA@15 by 20.3% on EgoHumans). Additionally, qualitative results show that our approach improves overall scene reconstruction quality. Our code is available at: https://github.com/hongsukchoi/HSfM_RELEASE

📄 PDF Abstract BibTeX arXiv:2412.17806

Code (1)

hongsukchoi/hsfm_release 공식 구현 pytorch

Tasks

Camera Pose EstimationPose Estimation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos

2026-06-01 · Jinpeng Liu, Yukang Xu, Yutong Li, Xingyu Liu arxiv

Reconstructing humans and their surrounding environments in a globally consistent 4D space is essential for comprehensive perception. However, prior works typically assume single-view inputs or decouple humans, scenes, a…

Spatial Reasoning

mID: Tracking and Identifying People with Millimeter Wave Radar

2019-05-29 · 2019 15th International Conference on Distributed Computing in Sensor Systems (DCOSS) 2019 5 · Peijun Zhao, Chris Xiaoxuan Lu, Jianan Wang, Changhao Chen 외

The key to offering personalised services in smart spaces is knowing where a particular person is with a high degree of accuracy. Visual tracking is one such solution, but concerns arise around the potential leakage of r…

RF-based Visual TrackingVisual Tracking

People Tracking in Panoramic Video for Guiding Robots

2022-06-06 · Alberto Bacchin, Filippo Berno, Emanuele Menegatti, Alberto Pretto

A guiding robot aims to effectively bring people to and from specific places within environments that are possibly unknown to them. During this operation the robot should be able to detect and track the accompanied perso…

Monitoring of people entering and exiting private areas using Computer Vision

2019-08-02 · Vinay Kumar V, P. Nagabhushan

Entry-Exit surveillance is a novel research problem that addresses security concerns when people attain absolute privacy in camera forbidden areas such as toilets and changing rooms that are basic amenities to the humans…

Event Detection

HUMBI: A Large Multiview Dataset of Human Body Expressions

2018-12-01 · CVPR 2020 6 · Zhixuan Yu, Jae Shin Yoon, In Kyu Lee, Prashanth Venkatesh 외

This paper presents a new large multiview dataset called HUMBI for human body expressions with natural clothing. The goal of HUMBI is to facilitate modeling view-specific appearance and geometry of gaze, face, hand, body…