paper-with-me

Papers

DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction

2026-03-03 · Yufu Wang, Evonne Ng, Soyong Shin, Rawal Khirodkar, Yuan Dong, Zhaoen Su, Jinhyung Park, Kris Kitani, Alexander Richard, Fabian Prada, Michael Zollhofer arxiv

We present DuoMo, a generative method that recovers human motion in world-space coordinates from unconstrained videos with noisy or incomplete observations. Reconstructing such motion requires solving a fundamental trade-off: generalizing from diverse and noisy video inputs while maintaining global motion consistency. Our approach addresses this problem by factorizing motion learning into two diffusion models. The camera-space model first estimates motion from videos in camera coordinates. The world-space model then lifts this initial estimate into world coordinates and refines it to be globally consistent. Together, the two models can reconstruct motion across diverse scenes and trajectories, even from highly noisy or incomplete observations. Moreover, our formulation is general, generating the motion of mesh vertices directly and bypassing parametric models. DuoMo achieves state-of-the-art performance. On EMDB, our method obtains a 16% reduction in world-space reconstruction error while maintaining low foot skating. On RICH, it obtains a 30% reduction in world-space error. Project page: https://yufu-wang.github.io/duomo/

📄 PDF Abstract BibTeX arXiv:2603.03265

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EmoScene: A Dual-space Dataset for Controllable Affective Image Generation

2026-04-01 · Li He, Longtai Zhang, Wenqiang Zhang, Yan Wang 외 arxiv

Text-to-image diffusion models have achieved high visual fidelity, yet precise control over scene semantics and fine-grained affective tone remains challenging. Human visual affect arises from the rapid integration of co…

Image Generation

VMC: Video Motion Customization using Temporal Attention Adaption for Text-to-Video Diffusion Models

2023-12-01 · CVPR 2024 1 · Hyeonho Jeong, Geon Yeong Park, Jong Chul Ye

Text-to-video diffusion models have advanced video generation significantly. However, customizing these models to generate videos with tailored motions presents a substantial challenge. In specific, they encounter hurdle…

Video EditingVideo Generation

Rethinking Video Super-Resolution: Towards Diffusion-Based Methods without Motion Alignment

2025-03-05 · Zhihao Zhan, Wang Pang, Xiang Zhu, Yechao Bai

In this work, we rethink the approach to video super-resolution by introducing a method based on the Diffusion Posterior Sampling framework, combined with an unconditional video diffusion transformer operating in latent …

AllSuper-ResolutionUnconditional Video GenerationVideo Generation+1

Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation

2026-01-21 · Yifei Liu, Changxing Ding, Ling Guo, Huaiguang Jiang 외 arxiv

Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impressive generative capabilities and flexibility. However, current motion diffusion models face two …

Two-in-One: Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer

2024-12-21 · Boyuan Li, Xihua Wang, Ruihua Song, Wenbing Huang

Multi-person interactive motion generation, a critical yet under-explored domain in computer character animation, poses significant challenges such as intricate modeling of inter-human interactions beyond individual moti…

Motion Generation