DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation
In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation. Despite their effectiveness, these frameworks face challenges in maintaining consistent narratives and handling shifts in scene composition or object placement from a single abstract user prompt. Exploring the ability of large language models (LLMs) to generate time-dependent, frame-by-frame prompts, this paper introduces a new framework, dubbed DirecT2V. DirecT2V leverages instruction-tuned LLMs as directors, enabling the inclusion of time-varying content and facilitating consistent video generation. To maintain temporal consistency and prevent mapping the value to a different object, we equip a diffusion model with a novel value mapping method and dual-softmax filtering, which do not require any additional training. The experimental results validate the effectiveness of our framework in producing visually coherent and storyful videos from abstract user prompts, successfully addressing the challenges of zero-shot video generation.
Code (1)
Tasks
Text-to-Video GenerationVideo GenerationZero-shot Text-to-Video GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
The hardcore brokers: Core-periphery structure and political representation in Denmark's corporate elite network
Who represents the corporate elite in democratic governance? Prior studies find a tightly integrated "inner circle" network representing the corporate elite politically across varieties of capitalism, yet they all rely o…
Evaluating the Effects of AI Directors for Quest Selection
Modern commercial games are designed for mass appeal, not for individual players, but there is a unique opportunity in video games to better fit the individual through adapting game elements. In this paper, we focus on A…
Cine-AI: Generating Video Game Cutscenes in the Style of Human Directors
Cutscenes form an integral part of many video games, but their creation is costly, time-consuming, and requires skills that many game developers lack. While AI has been leveraged to semi-automate cutscene production, the…
UnityVisualization of Board of Director Connections for Analysis in Socially Responsible Investing
This project is a collaboration between industry and academia to delve into Finance Social Networks, specifically the Board of Directors of public companies. Knowing the connections between Directors and Executives in di…
Data VisualizationCan an Agency Role-Reversal Lead to an Organizational Collapse?; A Study Proposal
The Principal-Agent Theory model is widely used to explain governance role where there is a separation of ownership and control, as it defines clear boundaries between governance and executives. However, examination of r…