paper-with-me

홈 › Papers

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

2026-09-24 · Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong hf

Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.

📄 PDF Abstract BibTeX arXiv:2609.30221

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video Generation

Similar Papers 제목 키워드 기반

Cut2Next: Generating Next Shot via In-Context Tuning

2025-08-11 · Jingwen He, Hongbo Liu, Jiajun Li, Ziqi Huang 외 arxiv

Effective multi-shot generation demands purposeful, film-like transitions and strict cinematic continuity. Current methods, however, often prioritize basic visual consistency, neglecting crucial editing patterns (e.g., s…

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

2026-07-27 · Shengyi Wang, Niantong Li, Guangzheng Hu, Hong Qi 외 hf

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained…

Video Generation

JAWS: Just A Wild Shot for Cinematic Transfer in Neural Radiance Fields

2023-03-27 · CVPR 2023 1 · Xi Wang, Robin Courant, Jinglei Shi, Eric Marchand 외

This paper presents JAWS, an optimization-driven approach that achieves the robust transfer of visual cinematic features from a reference in-the-wild video clip to a newly generated clip. To this end, we rely on an impli…

NeRF

ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation

2026-03-12 · Songlin Yang, Zhe Wang, Xuyi Yang, Songchun Zhang 외 arxiv

Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Implicit textual prompts lack precision, while explicit trajectory conditioni…

Video Generation

Customized Visual Storytelling with Unified Multimodal LLMs

2026-03-29 · Wei-Hua Li, Cheng Sun, Chu-Song Chen arxiv

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, …

Visual StorytellingStory Generation