paper-with-me

홈 › Papers

HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives

2025-10-23 · Yihao Meng, Hao Ouyang, Yue Yu, Qiuyu Wang, Wen Wang, Ka Leong Cheng, Hanlin Wang, Yixuan Li, Cheng Chen, Yanhong Zeng, Yujun Shen, Huamin Qu arxiv

State-of-the-art text-to-video models excel at generating isolated clips but fall short of creating the coherent, multi-shot narratives, which are the essence of storytelling. We bridge this "narrative gap" with HoloCine, a model that generates entire scenes holistically to ensure global consistency from the first shot to the last. Our architecture achieves precise directorial control through a Window Cross-Attention mechanism that localizes text prompts to specific shots, while a Sparse Inter-Shot Self-Attention pattern (dense within shots but sparse between them) ensures the efficiency required for minute-scale generation. Beyond setting a new state-of-the-art in narrative coherence, HoloCine develops remarkable emergent abilities: a persistent memory for characters and scenes, and an intuitive grasp of cinematic techniques. Our work marks a pivotal shift from clip synthesis towards automated filmmaking, making end-to-end cinematic creation a tangible future. Our code is available at: https://holo-cine.github.io/.

📄 PDF Abstract BibTeX arXiv:2510.20822

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ShotPlan: Cinematic Video Generation with Learnable Planning Token

2026-07-20 · Su Guo, Guangce Liu, Haosen Yang, Jiepeng Wang 외 hf

Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition require explicit shot…

Video Generation

Can video generation replace cinematographers? Research on the cinematic language of generated video

2024-12-16 · Xiaozhe Li, Kai Wu, Siyi Yang, YiZhan Qu 외

Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing research primarily focuses on object mo…

Video Generation

Camera Artist: A Multi-Agent Framework for Cinematic Language Storytelling Video Generation

2026-04-10 · Haobo Hu, Qi Mao, Yuanhang Li, Libiao Jin arxiv

We propose Camera Artist, a multi-agent framework that models a real-world filmmaking workflow to generate narrative videos with explicit cinematic language. While recent multi-agent systems have made substantial progres…

Video Generation

STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative

2025-12-13 · Peixuan Zhang, Zijian Jia, Kaiqi Liu, Shuchen Weng 외 arxiv

While recent advancements in generative models have achieved remarkable visual fidelity in video synthesis, creating coherent multi-shot narratives remains a significant challenge. To address this, keyframe-based approac…

Video Generation

ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation

2026-03-12 · Songlin Yang, Zhe Wang, Xuyi Yang, Songchun Zhang 외 arxiv

Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Implicit textual prompts lack precision, while explicit trajectory conditioni…

Video Generation