paper-with-me

Papers

ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation

2026-03-12 · Songlin Yang, Zhe Wang, Xuyi Yang, Songchun Zhang, Xianghao Kong, Taiyi Wu, Xiaotong Zhao, Ran Zhang, Alan Zhao, Anyi Rao arxiv

Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Implicit textual prompts lack precision, while explicit trajectory conditioning imposes prohibitive manual overhead and often triggers execution failures in current models. To overcome this bottleneck, we propose a data-centric paradigm shift, positing that aligned (Caption, Trajectory, Video) triplets form an inherent joint distribution that can connect automated plotting and precise execution. Guided by this insight, we present ShotVerse, a ``Plan-then-Control'' framework that decouples generation into two collaborative agents: a VLM (Vision-Language Model)-based Planner that leverages spatial priors to obtain cinematic, globally aligned trajectories from text, and a Controller that renders these trajectories into multi-shot video content via a camera adapter. Central to our approach is the construction of a data foundation: we design an automated multi-shot camera calibration pipeline aligns disjoint single-shot trajectories into a unified global coordinate system. This facilitates the curation of ShotVerse-Bench, a high-fidelity cinematic dataset with a three-track evaluation protocol that serves as the bedrock for our framework. Extensive experiments demonstrate that ShotVerse effectively bridges the gap between unreliable textual control and labor-intensive manual plotting, achieving superior cinematic aesthetics and generating multi-shot videos that are both camera-accurate and cross-shot consistent.

📄 PDF Abstract BibTeX arXiv:2603.11421

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

Can video generation replace cinematographers? Research on the cinematic language of generated video

2024-12-16 · Xiaozhe Li, Kai Wu, Siyi Yang, YiZhan Qu 외

Recent advancements in text-to-video (T2V) generation have leveraged diffusion models to enhance visual coherence in videos synthesized from textual descriptions. However, existing research primarily focuses on object mo…

Video Generation

CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation

2026-02-06 · Kaiyi Huang, Yukun Huang, Yu Li, Jianhong Bai 외 arxiv

Cinematic video production requires control over scene-subject composition and camera movement, but live-action shooting remains costly due to the need for constructing physical sets. To address this, we introduce the ta…

Text-to-Video Generation

CineLOG: A Training Free Approach for Cinematic Long Video Generation

2025-12-13 · Zahra Dehghanian, Morteza Abolghasemi, Hamid Beigy, Hamid R. Rabiee arxiv

Controllable video synthesis is a central challenge in computer vision, yet current models struggle with fine grained control beyond textual prompts, particularly for cinematic attributes like camera trajectory and genre…

Video Generation

Generative Photographic Control for Scene-Consistent Video Cinematic Editing

2025-11-17 · Huiqiang Sun, Liao Shen, Zhan Peng, Kun Wang 외 arxiv

Cinematic storytelling is profoundly shaped by the artful manipulation of photographic elements such as depth of field and exposure. These effects are crucial in conveying mood and creating aesthetic appeal. However, con…

AKiRa: Augmentation Kit on Rays for optical video generation

2024-12-18 · CVPR 2025 1 · Xi Wang, Robin Courant, Marc Christie, Vicky Kalogeiton

Recent advances in text-conditioned video diffusion have greatly improved video quality. However, these methods offer limited or sometimes no control to users on camera aspects, including dynamic camera motion, zoom, dis…

Video Generation