paper-with-me

홈 › Papers

Towards Geometry-Aware and Motion-Guided Video Human Mesh Recovery

2026-01-29 · Hongjun Chen, Huan Zheng, Wencheng Han, Jianbing Shen arxiv

Existing video-based 3D Human Mesh Recovery (HMR) methods often produce physically implausible results, stemming from their reliance on flawed intermediate 3D pose anchors and their inability to effectively model complex spatiotemporal dynamics. To overcome these deep-rooted architectural problems, we introduce HMRMamba, a new paradigm for HMR that pioneers the use of Structured State Space Models (SSMs) for their efficiency and long-range modeling prowess. Our framework is distinguished by two core contributions. First, the Geometry-Aware Lifting Module, featuring a novel dual-scan Mamba architecture, creates a robust foundation for reconstruction. It directly grounds the 2D-to-3D pose lifting process with geometric cues from image features, producing a highly reliable 3D pose sequence that serves as a stable anchor. Second, the Motion-guided Reconstruction Network leverages this anchor to explicitly process kinematic patterns over time. By injecting this crucial temporal awareness, it significantly enhances the final mesh's coherence and robustness, particularly under occlusion and motion blur. Comprehensive evaluations on 3DPW, MPI-INF-3DHP, and Human3.6M benchmarks confirm that HMRMamba sets a new state-of-the-art, outperforming existing methods in both reconstruction accuracy and temporal consistency while offering superior computational efficiency.

📄 PDF Abstract BibTeX arXiv:2601.21376

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyLong-range modelingHuman Mesh Recovery

Similar Papers 제목 키워드 기반

CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives

2025-12-16 · Zihan Wang, Jiashun Wang, Jeff Tan, Yiwen Zhao 외 arxiv

We introduce CRISP, a method that recovers simulatable human motion and scene geometry from monocular video. Prior work on joint human-scene reconstruction relies on data-driven priors and joint optimization with no phys…

Reinforcement Learning

Geometry-Guided Camera Motion Understanding in VideoLLMs

2026-03-13 · Haoan Feng, Sri Harsha Musunuri, Guan-Ming Su arxiv

Camera motion is a fundamental geometric signal that shapes visual perception and cinematic style, yet current video-capable vision-language models (VideoLLMs) rarely represent it explicitly and often fail on fine-graine…

Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation

2025-12-04 · Yanran Zhang, Ziyi Wang, Wenzhao Zheng, Zheng Zhu 외 arxiv

Generating interactive and dynamic 4D scenes from a single static image remains a core challenge. Most existing generate-then-reconstruct and reconstruct-then-generate methods decouple geometry from motion, causing spati…

PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence

2025-12-15 · Ruiyan Wang, Teng Hu, Kaihui Huang, Zihan Su 외 arxiv

Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. Howeve…

Video Generation

DeCo: Decoupled Human-Centered Diffusion Video Editing with Motion Consistency

2024-08-14 · Xiaojing Zhong, Xinyi Huang, Xiaofeng Yang, Guosheng Lin 외

Diffusion models usher a new era of video editing, flexibly manipulating the video contents with text prompts. Despite the widespread application demand in editing human-centered videos, these models face significant cha…

text-guided-image-editingVideo Editing