paper-with-me

Papers

Human3R: Everyone Everywhere All at Once

2025-10-07 · Yue Chen, Xingyu Chen, Yuxuan Xue, Anpei Chen, Yuliang Xiu, Gerard Pons-Moll arxiv

We present Human3R, a unified, feed-forward framework for online 4D human-scene reconstruction, in the world frame, from casually captured monocular videos. Unlike previous approaches that rely on multi-stage pipelines, iterative contact-aware refinement between humans and scenes, and heavy dependencies, e.g., human detection, depth estimation, and SLAM pre-processing, Human3R jointly recovers global multi-person SMPL-X bodies ("everyone"), dense 3D scene ("everywhere"), and camera trajectories in a single forward pass ("all-at-once"). Our method builds upon the 4D online reconstruction model CUT3R, and uses parameter-efficient visual prompt tuning, to strive to preserve CUT3R's rich spatiotemporal priors, while enabling direct readout of multiple SMPL-X bodies. Human3R is a unified model that eliminates heavy dependencies and iterative refinement. After being trained on the relatively small-scale synthetic dataset BEDLAM for just one day on one GPU, it achieves superior performance with remarkable efficiency: it reconstructs multiple humans in a one-shot manner, along with 3D scenes, in one stage, in real-time (15 FPS) with a low memory footprint (8 GB). Extensive experiments demonstrate that Human3R delivers state-of-the-art or competitive performance across tasks, including global human motion estimation, local human mesh recovery, video depth estimation, and camera pose estimation, with a single unified model. We hope that Human3R will serve as a simple yet strong baseline, which can be easily adapted for downstream applications. Code, models and 4D interactive demos are available at https://fanegg.github.io/Human3R/.

📄 PDF Abstract BibTeX arXiv:2510.06219

Code (0)

등록된 구현이 없습니다.

Tasks

Camera Pose EstimationVisual Prompt TuningHuman Mesh RecoveryDepth Estimation

Similar Papers 제목 키워드 기반

GUSH3R: Everyone Everywhere All at Once as Gaussians

2026-07-06 · Keito Abe, Kaede Shiohara, Takashi Otonari, Toshihiko Yamasaki arxiv

Reconstructing dynamic human-scene environments from monocular videos is a challenging problem that requires jointly modeling scene geometry, camera motion, and non-rigid human dynamics while enabling photorealistic rend…

Novel View Synthesis3D ReconstructionPoint Clouds

Mocap Everyone Everywhere: Lightweight Motion Capture With Smartwatches and a Head-Mounted Camera

2024-01-01 · CVPR 2024 1 · Jiye Lee, Hanbyul Joo

We present a lightweight and affordable motion capture method based on two smartwatches and a head-mounted camera. In contrast to the existing approaches that use six or more expert-level IMU devices, our approach is muc…

Motion Estimation

Personalized and Demand-Based Education Concept: Practical Tools for Control Engineers

2025-04-10 · Balint Varga, Lars Fischer, Levente Kovacs

This paper presents a personalized lecture concept using educational blocks and its demonstrative application in a new university lecture. Higher education faces daily challenges: deep and specialized knowledge is availa…

SWift -- A SignWriting improved fast transcriber

2019-11-25 · Claudia S. Bianchini, Fabrizio Borgia, Paolo Bottoni, Maria de Marsico

We present SWift (SignWriting improved fast transcriber), an advanced editor for computer-aided writing and transcribing using SignWriting (SW). SW is devised to allow deaf people and linguists alike to exploit an easy-t…

Everywhere Learning: Artificial Intelligence with Pointwise Constraints

2026-06-01 · Ignacio Boero, Ignacio Hounie, Luiz Chamon, Alejandro Ribeiro arxiv

Everywhere learning is a new paradigm whereby Artificial Intelligence (AI) systems are trained to satisfy loss constraints with probability one over the data distribution. This is in contrast to the standard paradigm of …