paper-with-me

홈 › Papers

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos

2026-01-15 · Chengfeng Zhao, Jiazhi Shu, Yubo Zhao, Tianyu Huang, Jiahao Lu, Zekai Gu, Chengwei Ren, Zhiyang Dou, Qing Shuai, Yuan Liu arxiv

In this paper, we find that the generation of 3D human motions and 2D human videos is intrinsically coupled. 3D motions provide the structural prior for plausibility and consistency in videos, while pre-trained video models offer strong generalization capabilities for motions. Based on this, we present CoMoVi, a co-generative framework that generates 3D human motions and videos synchronously within a single diffusion denoising loop. However, since the 3D human motions and the 2D human-centric videos have a modality gap between each other, we propose to project the 3D human motion into an effective 2D human motion representation that effectively aligns with the 2D videos. Then, we design a dual-branch diffusion model to couple human motion and the video generation process with mutual feature interaction and 3D-2D cross attentions. To train and evaluate our model, we curate CoMoVi-Dataset, a large-scale real-world human video dataset with text and motion annotations, covering diverse and challenging human motions. Extensive experiments demonstrate that our method generates high-quality 3D human motion with a better generalization ability and that our method can generate high-quality human-centric videos without external motion references.

📄 PDF Abstract BibTeX arXiv:2601.10632

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

ViMo: Generating Motions from Casual Videos

2024-08-13 · Liangdong Qiu, Chengxing Yu, Yanran Li, Zhao Wang 외

Although humans have the innate ability to imagine multiple possible actions from videos, it remains an extraordinary challenge for computers due to the intricate camera movements and montages. Most existing motion gener…

Motion Generation

Object-Aware 4D Human Motion Generation

2025-10-31 · Shurui Gui, Deep Anil Patel, Xiner Li, Martin Renqiang Min arxiv

Recent advances in video diffusion models have enabled the generation of high-quality videos. However, these videos still suffer from unrealistic deformations, semantic violations, and physical inconsistencies that are l…

Emotional Talking Faces: Making Videos More Expressive and Realistic

2022-12-13 · ACM Multimedia Asia 2022 12 · Sahil Goyal, Shagun Uppal, Sarthak Bhagat, Dhroov Goel 외

Lip synchronization and talking face generation have gained a specific interest from the research community with the advent and need of digital communication in different fields. Prior works propose several elegant solut…

Face GenerationTalking Face Generation

HumanScore: Benchmarking Human Motions in Generated Videos

2026-04-22 · Yusu Fang, Tiange Xiang, Tian Tan, Narayan Schuetz 외 arxiv

Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, no prior method systematically measures how faithfully these …

Video Generation

Emotionally Enhanced Talking Face Generation

2023-03-21 · Sahil Goyal, Shagun Uppal, Sarthak Bhagat, Yi Yu 외

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to crea…

Face GenerationTalking Face GenerationTalking Head Generation