paper-with-me

홈 › Papers

WHAM: Reconstructing World-grounded Humans with Accurate 3D Motion

2023-12-12 · CVPR 2024 1 · Soyong Shin, Juyong Kim, Eni Halilaj, Michael J. Black

The estimation of 3D human motion from video has progressed rapidly but current methods still have several key limitations. First, most methods estimate the human in camera coordinates. Second, prior work on estimating humans in global coordinates often assumes a flat ground plane and produces foot sliding. Third, the most accurate methods rely on computationally expensive optimization pipelines, limiting their use to offline applications. Finally, existing video-based methods are surprisingly less accurate than single-frame methods. We address these limitations with WHAM (World-grounded Humans with Accurate Motion), which accurately and efficiently reconstructs 3D human motion in a global coordinate system from video. WHAM learns to lift 2D keypoint sequences to 3D using motion capture data and fuses this with video features, integrating motion context and visual information. WHAM exploits camera angular velocity estimated from a SLAM method together with human motion to estimate the body's global trajectory. We combine this with a contact-aware trajectory refinement method that lets WHAM capture human motion in diverse conditions, such as climbing stairs. WHAM outperforms all existing 3D human motion recovery methods across multiple in-the-wild benchmarks. Code will be available for research purposes at http://wham.is.tue.mpg.de/

📄 PDF Abstract BibTeX arXiv:2312.07531

Code (1)

yohanshin/WHAM pytorch

Tasks

3D Human Pose Estimation

Similar Papers 제목 키워드 기반

WhAM: Towards A Translative Model of Sperm Whale Vocalization

2025-12-01 · Orr Paradise, Pranav Muralikrishnan, Liangyuan Chen, Hugo Flores García 외 arxiv

Sperm whales communicate in short sequences of clicks known as codas. We present WhAM (Whale Acoustics Model), the first transformer-based model capable of generating synthetic sperm whale codas from any audio prompt. Wh…

AHAP: Reconstructing Arbitrary Humans from Arbitrary Perspectives with Geometric Priors

2026-02-27 · Xiaozhen Qiao, Wenjia Wang, Zhiyuan Zhao, Jiacheng Sun 외 arxiv

Reconstructing 3D humans from images captured at multiple perspectives typically requires pre-calibration, like using checkerboards or MVS algorithms, which limits scalability and applicability in diverse real-world scen…

Camera Pose EstimationContrastive Learning

Reconstructing People, Places, and Cameras

2024-12-23 · CVPR 2025 1 · Lea Müller, Hongsuk Choi, Anthony Zhang, Brent Yi 외

We present "Humans and Structure from Motion" (HSfM), a method for jointly reconstructing multiple human meshes, scene point clouds, and camera parameters in a metric world coordinate system from a sparse set of uncalibr…

Camera Pose EstimationPose Estimation

ALARM: Active LeArning of Rowhammer Mitigations

2022-11-30 · Amir Naseredini, Martin Berger, Matteo Sammartino, Shale Xiong

Rowhammer is a serious security problem of contemporary dynamic random-access memory (DRAM) where reads or writes of bits can flip other bits. DRAM manufacturers add mitigations, but don't disclose details, making it dif…

Active Learning

WHAC: World-grounded Humans and Cameras

2024-03-19 · Wanqi Yin, Zhongang Cai, Ruisi Wang, Fanzhou Wang 외

Estimating human and camera trajectories with accurate scale in the world coordinate system from a monocular video is a highly desirable yet challenging and ill-posed problem. In this study, we aim to recover expressive …

Camera Pose EstimationPose Estimation