paper-with-me

홈 › Papers

TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-Animation

2026-06-10 · Cheng-Feng Pu, Jia-Peng Zhang, Meng-Hao Guo, Yan-Pei Cao, Shi-Min Hu arxiv

The explosion of generative 3D assets has created a massive demand for animation, yet current motion capture methods remain brittle, restricted to species-specific templates (e.g., SMPL) or requiring labor-intensive manual rigging. We introduce TopoCap, the first unified framework capable of extracting motion from monocular video and retargeting it onto characters with arbitrary, unseen skeletal topologies, i.e., from bipeds to hexapods and inanimate objects, without test-time optimization. Our key insight is that while skeletal structures are combinatorial and discrete, the underlying physics of motion occupy a continuous, low-dimensional manifold. We materialize this insight via a two-stage generative pipeline. First, we learn a Universal Motion Manifold using a Graph CVAE that compresses heterogeneous kinematic chains into a shared, fixed-length latent code. By explicitly conditioning the decoder on a structural embedding of the target rig, we disentangle motion dynamics from skeletal topology. Second, we treat video-to-animation as a conditional flow matching problem, predicting these topology-agnostic codes from visual features. To learn this generalized prior, we introduce Mobjaverse, a massive-scale dataset curated from Objaverse-XL. Comprising over 5,000 unique skeletal topologies and 2 million frames, it exceeds the structural diversity of existing datasets by two orders of magnitude. Extensive experiments demonstrate that \MethodMotion outperforms specialist models on human and quadruped benchmarks while enabling zero-shot retargeting for the long tail of 3D creatures. Dataset is publicly available at https://huggingface.co/datasets/duckduckplz/Mobjaverse.

📄 PDF Abstract BibTeX arXiv:2606.12153

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Human Motion from Monocular Videos via Cross-Modal Manifold Alignment

2024-04-15 · Shuaiying Hou, Hongyu Tao, Junheng Fang, Changqing Zou 외

Learning 3D human motion from 2D inputs is a fundamental task in the realms of computer vision and computer graphics. Many previous methods grapple with this inherently ambiguous task by introducing motion priors into th…

NECromancer: Breathing Life into Skeletons via BVH Animation

2026-02-06 · Mingxi Xu, Qi Wang, Zhengyu Wen, Phong Dao Thien 외 arxiv

Motion tokenization is a key component of generalizable motion models, yet most existing approaches are restricted to species-specific skeletons, limiting their applicability across diverse morphologies. We propose NECro…

MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion

2025-04-28 · CVPR 2025 1 · Zador Pataki, Paul-Edouard Sarlin, Johannes L. Schönberger, Marc Pollefeys

While Structure-from-Motion (SfM) has seen much progress over the years, state-of-the-art systems are prone to failure when facing extreme viewpoint changes in low-overlap, low-parallax or high-symmetry scenarios. Becaus…

One Video, One World: Turning Monocular Video into Physical 4D Scenes

2026-06-30 · Junhao Chen, Boran Zhang, Mingjin Chen, Henghaofan Zhang 외 arxiv

We introduce \textbf{OVOW}, the first training-free system that reconstructs \emph{instance-level, simulation-ready} 4D mesh scenes from a single monocular video. Recent 4D reconstruction achieves impressive rendering qu…

Point Clouds

4D Monocular Surgical Reconstruction under Arbitrary Camera Motions

2026-02-19 · Jiwei Shan, Zeyu Cai, Cheng-Tai Hsieh, Yirui Li 외 arxiv

Reconstructing deformable surgical scenes from endoscopic videos is challenging and clinically important. Recent state-of-the-art methods based on implicit neural representations or 3D Gaussian splatting have made notabl…