paper-with-me

홈 › Papers

Masked Modeling for Human Motion Recovery Under Occlusions

2026-01-22 · Zhiyin Qian, Siwei Zhang, Bharat Lal Bhatnagar, Federica Bogo, Siyu Tang arxiv

Human motion reconstruction from monocular videos is a fundamental challenge in computer vision, with broad applications in AR/VR, robotics, and digital content creation, but remains challenging under frequent occlusions in real-world settings. Existing regression-based methods are efficient but fragile to missing observations, while optimization- and diffusion-based approaches improve robustness at the cost of slow inference speed and heavy preprocessing steps. To address these limitations, we leverage recent advances in generative masked modeling and present MoRo: Masked Modeling for human motion Recovery under Occlusions. MoRo is an occlusion-robust, end-to-end generative framework that formulates motion reconstruction as a video-conditioned task, and efficiently recover human motion in a consistent global coordinate system from RGB videos. By masked modeling, MoRo naturally handles occlusions while enabling efficient, end-to-end inference. To overcome the scarcity of paired video-motion data, we design a cross-modality learning scheme that learns multi-modal priors from a set of heterogeneous datasets: (i) a trajectory-aware motion prior trained on MoCap datasets, (ii) an image-conditioned pose prior trained on image-pose datasets, capturing diverse per-frame poses, and (iii) a video-conditioned masked transformer that fuses motion and pose priors, finetuned on video-motion datasets to integrate visual cues with motion dynamics for robust inference. Extensive experiments on EgoBody and RICH demonstrate that MoRo substantially outperforms state-of-the-art methods in accuracy and motion realism under occlusions, while performing on-par in non-occluded scenarios. MoRo achieves real-time inference at 70 FPS on a single H200 GPU.

📄 PDF Abstract BibTeX arXiv:2601.16079

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DiMo: Discrete Diffusion Modeling for Motion Generation and Understanding

2026-02-04 · Ning Zhang, Zhengyu Li, Kwong Weng Loh, Mingxi Xu 외 arxiv

Prior masked modeling motion generation methods predominantly study text-to-motion. We present DiMo, a discrete diffusion-style framework, which extends masked modeling to bidirectional text--motion understanding and gen…

MoMask: Generative Masked Modeling of 3D Human Motions

2023-11-29 · CVPR 2024 1 · Chuan Guo, Yuxuan Mu, Muhammad Gohar Javed, Sen Wang 외

We introduce MoMask, a novel masked modeling framework for text-driven 3D human motion generation. In MoMask, a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens…

Human motion predictionMotion ForecastingMotion GenerationMotion Interpolation+1

Leveraging MoCap Data for Human Mesh Recovery

2021-10-18 · Fabien Baradel, Thibault Groueix, Philippe Weinzaepfel, Romain Brégier 외

Training state-of-the-art models for human body pose and shape recovery from images or videos requires datasets with corresponding annotations that are really hard and expensive to obtain. Our goal in this paper is to st…

3D Human Pose Estimation3D Human Reconstruction3D Human Shape EstimationHuman Mesh Recovery

Text-driven Human Motion Generation with Motion Masked Diffusion Model

2024-09-29 · Xingyu Chen

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, wh…

DiversityMotion Generation

Masked Motion Predictors are Strong 3D Action Representation Learners

2023-08-14 · ICCV 2023 1 · Yunyao Mao, Jiajun Deng, Wengang Zhou, Yao Fang 외

In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating e…

3D Action RecognitionAction RecognitionFew-Shot Skeleton-Based Action Recognitionmotion prediction+3