paper-with-me

홈 › Papers

UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework

2025-12-03 · Youxin Pang, Yong Zhang, Ruizhi Shao, Xiang Deng, Feng Gao, Xu Xiaoming, Xiaoming Wei, Yebin Liu arxiv

We propose UniMo, an innovative autoregressive model for joint modeling of 2D human videos and 3D human motions within a unified framework, enabling simultaneous generation and understanding of these two modalities for the first time. Current methods predominantly focus on generating one modality given another as the condition or integrating either of them with other modalities such as text and audio. Unifying 2D videos and 3D motions for simultaneous optimization and generation remains largely unexplored, presenting significant challenges due to their substantial structural and distributional differences. Inspired by the LLM's ability to unify different modalities, our method models videos and 3D motions as a unified tokens sequence, utilizing separate embedding layers to mitigate distribution gaps. Additionally, we devise a sequence modeling strategy that integrates two distinct tasks within a single framework, proving the effectiveness of unified modeling. Moreover, to efficiently align with visual tokens and preserve 3D spatial information, we design a novel 3D motion tokenizer with a temporal expansion strategy, using a single VQ-VAE to produce quantized motion tokens. It features multiple expert decoders that handle body shapes, translation, global orientation, and body poses for reliable 3D motion reconstruction. Extensive experiments demonstrate that our method simultaneously generates corresponding videos and motions while performing accurate motion capture. This work taps into the capacity of LLMs to fuse diverse data types, paving the way for integrating human-centric information into existing models and potentially enabling multimodal, controllable joint modeling of humans, objects, and scenes.

📄 PDF Abstract BibTeX arXiv:2512.03918

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unimotion: Unifying 3D Human Motion Synthesis and Understanding

2024-09-24 · Chuqiao Li, Julian Chibane, Yannan He, Naama Pearl 외

We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control avatar motion with global text conditioni…

Motion Synthesis

BAMM: Bidirectional Autoregressive Motion Model

2024-03-28 · Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Pu Wang, Minwoo Lee 외

Generating human motion from text has been dominated by denoising motion models either through diffusion or generative masking process. However, these models face great limitations in usability by requiring prior knowled…

DenoisingmodelMotion GenerationMotion Synthesis

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

2026-06-26 · Haoyuan Wang, Yabo Chen, Haibin Huang, Chi Zhang 외 arxiv

Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video generation offers a scalable foundation, but suffers from error accu…

Video Generation

X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents

2025-08-12 · Guoxian Song, Hongyi Xu, Xiaochen Zhao, You Xie 외 arxiv

We present X-UniMotion, a unified and expressive implicit latent representation for whole-body human motion, encompassing facial expressions, body poses, and hand gestures. Unlike prior motion transfer methods that rely …

Autoregressive Flow Matching for Motion Prediction

2025-12-27 · Johnathan Xie, Stefan Stojanov, Cristobal Eyzaguirre, Daniel L. K. Yamins 외 arxiv

Motion prediction has been studied in different contexts with models trained on narrow distributions and applied to downstream tasks in human motion prediction and robotics. Simultaneously, recent efforts in scaling vide…

Video GenerationVideo Prediction