paper-with-me

Papers

MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos

2025-12-11 · Kehong Gong, Zhengyu Wen, Weixia He, Mingxi Xu, Qi Wang, Ning Zhang, Zhengyu Li, Dongze Lian, Wei Zhao, Xiaoyu He, Mingyuan Zhang arxiv

Motion capture now underpins content creation far beyond digital humans, yet most existing pipelines remain species- or template-specific. We formalize this gap as Category-Agnostic Motion Capture (CAMoCap): given a monocular video and an arbitrary rigged 3D asset as a prompt, the goal is to reconstruct a rotation-based animation such as BVH that directly drives the specific asset. We present MoCapAnything, a reference-guided, factorized framework that first predicts 3D joint trajectories and then recovers asset-specific rotations via constraint-aware inverse kinematics. The system contains three learnable modules and a lightweight IK stage: (1) a Reference Prompt Encoder that extracts per-joint queries from the asset's skeleton, mesh, and rendered images; (2) a Video Feature Extractor that computes dense visual descriptors and reconstructs a coarse 4D deforming mesh to bridge the gap between video and joint space; and (3) a Unified Motion Decoder that fuses these cues to produce temporally coherent trajectories. We also curate Truebones Zoo with 1038 motion clips, each providing a standardized skeleton-mesh-render triad. Experiments on both in-domain benchmarks and in-the-wild videos show that MoCapAnything delivers high-quality skeletal animations and exhibits meaningful cross-species retargeting across heterogeneous rigs, enabling scalable, prompt-driven 3D motion capture for arbitrary assets. Project page: https://animotionlab.github.io/MoCapAnything/

📄 PDF Abstract BibTeX arXiv:2512.10881

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MoCapAnything V2: End-to-End Motion Capture for Arbitrary Skeletons

2026-04-30 · Kehong Gong, Zhengyu Wen, Dao Thien Phong, Mingxi Xu 외 arxiv

Recent methods for arbitrary-skeleton motion capture from monocular video follow a factorized pipeline, where a Video-to-Pose network predicts joint positions and an analytical inverse-kinematics (IK) stage recovers join…

NECromancer: Breathing Life into Skeletons via BVH Animation

2026-02-06 · Mingxi Xu, Qi Wang, Zhengyu Wen, Phong Dao Thien 외 arxiv

Motion tokenization is a key component of generalizable motion models, yet most existing approaches are restricted to species-specific skeletons, limiting their applicability across diverse morphologies. We propose NECro…

UniMate: One Unified Model to Animate Diverse Skeletons

2026-09-04 · Linzhan Mou, Jiahui Lei, Zhiyang Dou, Chenyue Cai 외 hf

Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion to drive them remains a bottleneck. Existing learned animators are topology-constrained: they rely on categor…

UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons

2023-09-13 · Sicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li 외

The automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability acros…

DiversityGesture Generation

AnyTop: Character Animation Diffusion with Any Topology

2025-02-24 · Inbar Gat, Sigal Raab, Guy Tevet, Yuval Reshef 외

Generating motion for arbitrary skeletons is a longstanding challenge in computer graphics, remaining largely unexplored due to the scarcity of diverse datasets and the irregular nature of the data. In this work, we intr…

Denoising