paper-with-me

홈 › Papers

DirectorLLM for Human-Centric Video Generation

2024-12-19 · Kunpeng Song, Tingbo Hou, Zecheng He, Haoyu Ma, Jialiang Wang, Animesh Sinha, Sam Tsai, Yaqiao Luo, Xiaoliang Dai, Li Chen, Xide Xia, Peizhao Zhang, Peter Vajda, Ahmed Elgammal, Felix Juefei-Xu

In this paper, we introduce DirectorLLM, a novel video generation model that employs a large language model (LLM) to orchestrate human poses within videos. As foundational text-to-video models rapidly evolve, the demand for high-quality human motion and interaction grows. To address this need and enhance the authenticity of human motions, we extend the LLM from a text generator to a video director and human motion simulator. Utilizing open-source resources from Llama 3, we train the DirectorLLM to generate detailed instructional signals, such as human poses, to guide video generation. This approach offloads the simulation of human motion from the video generator to the LLM, effectively creating informative outlines for human-centric scenes. These signals are used as conditions by the video renderer, facilitating more realistic and prompt-following video generation. As an independent LLM module, it can be applied to different video renderers, including UNet and DiT, with minimal effort. Experiments on automatic evaluation benchmarks and human evaluations show that our model outperforms existing ones in generating videos with higher human motion fidelity, improved prompt faithfulness, and enhanced rendered subject naturalness.

📄 PDF Abstract BibTeX arXiv:2412.14484

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelVideo Generation

Methods 이 논문이 사용한 방법론

LLaMA LLaMA is a collection of foundation language models ranging from 7B to 65B parameters. It is based on the transformer architecture with various improvements that were…

Similar Papers 제목 키워드 기반

OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

2024-11-28 · CVPR 2025 1 · Hui Li, Mingwang Xu, Yun Zhan, Shan Mu 외

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of…

Video Generation

EgoReAct: Egocentric Video-Driven 3D Human Reaction Generation

2025-12-28 · Libo Zhang, Zekun Li, Tianyu Li, Zeyu Cao 외 arxiv

Humans exhibit adaptive, context-sensitive responses to egocentric visual input. However, faithfully modeling such reactions from egocentric video remains challenging due to the dual requirements of strictly causal gener…

StreamingEffect: Real-Time Human-Centric Video Effect Generation

2026-05-16 · Yiren Song, Cheng Liu, Yuxin Jiang, Mike Zheng Shou arxiv

Streaming video effect generation is highly desirable for live human-centric applications such as e-commerce streaming, entertainment, and vlogging, yet remains difficult due to the lack of suitable data and deployable e…

Text-to-Video Generation

Intention-driven Ego-to-Exo Video Generation

2024-03-14 · Hongchen Luo, Kai Zhu, Wei Zhai, Yang Cao

Ego-to-exo video generation refers to generating the corresponding exocentric video according to the egocentric video, providing valuable applications in AR/VR and embodied AI. Benefiting from advancements in diffusion m…

Optical Flow EstimationStereo MatchingVideo Generation

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

2024-11-13 · XiaoFeng Wang, Kang Zhao, Feng Liu, Jiayu Wang 외

Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective…

Video Generation