paper-with-me

홈 › Papers

Towards 3D-Aware Video Diffusion Models: Render-Free Human Motion Control with Mesh Tokenization

2026-06-01 · Jingyun Liang, Min Wei, Shikai Li, Yizeng Han, Hangjie Yuan, Lei Sun, Weihua Chen, Fan Wang arxiv

Diffusion models have shown remarkable success in video generation. However, whether such models are truly aware of the 3D structure underlying visual observations, rather than simply reproducing plausible 2D projections, remains an open question. In this work, we investigate this question through human motion control, a task that requires precise modelling of 3D human geometry, motion, camera viewpoint, and scene context. Unlike prior methods that rely on rendered 2D motion guidance videos, we propose a render-free framework that conditions video generation directly on compressed 3D human mesh tokens. This representation preserves full 3D geometric information while enabling a unified token-based generation pipeline that processes video tokens jointly with motion tokens in a DiT-based architecture. This design requires the model to reason jointly about appearance, 3D structure, and camera viewpoint during video generation. Experimental results demonstrate strong performance on human motion control benchmarks, while reducing artifacts induced by view-dependent 2D guidance and trajectory-pose mismatches during editing. These findings suggest that video diffusion models, when equipped with mesh tokenization, can better capture complex 3D human structures and their interactions with the surrounding environment.

📄 PDF Abstract BibTeX arXiv:2606.02000

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

AMG: Avatar Motion Guided Video Generation

2024-09-02 · Zhangsihao Yang, Mengyi Shan, Mohammad Farazi, Wenhui Zhu 외

Human video generation task has gained significant attention with the advancement of deep generative models. Generating realistic videos with human movements is challenging in nature, due to the intricacies of human body…

Video Generation

From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection

2026-08-12 · Zepeng Wang, Jiagao Hu, Fuhao Li, Yuxuan Chen 외 hf

Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although single-image reflection removal has been extensively studied, video reflection remo…

Reflection Removal

iButter: Neural Interactive Bullet Time Generator for Human Free-viewpoint Rendering

2021-08-12 · Liao Wang, Ziyu Wang, Pei Lin, Yuheng Jiang 외

Generating ``bullet-time'' effects of human free-viewpoint videos is critical for immersive visual effects and VR/AR experience. Recent neural advances still lack the controllable and interactive bullet-time design abili…

NeRFVideo Generation

4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans

2026-07-30 · Renlong Wu, Haoran Chen, Yuxiang Wei, Xiaowei Jin 외 arxiv

Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D representation, which is expensive and …

DANCER: Dance ANimation via Condition Enhancement and Rendering with diffusion model

2025-10-31 · Yucheng Xing, Jinxing Yin, Xiaodong Liu arxiv

Recently, diffusion models have shown their impressive ability in visual generation tasks. Besides static images, more and more research attentions have been drawn to the generation of realistic videos. The video generat…

Video Generation