paper-with-me

홈 › Papers

CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation

2025-05-21 · Xinran Wang, Songyu Xu, Xiangxuan Shan, Yuxuan Zhang, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, Zhanyu Ma

Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of current models to grasp and reproduce cinematographic techniques remains largely uncharted, hindered by the scarcity of expert-annotated data. To bridge this gap, we present CineTechBench, a pioneering benchmark founded on precise, manual annotation by seasoned cinematography experts across key cinematography dimensions. Our benchmark covers seven essential aspects-shot scale, shot angle, composition, camera movement, lighting, color, and focal length-and includes over 600 annotated movie images and 120 movie clips with clear cinematographic techniques. For the understanding task, we design question answer pairs and annotated descriptions to assess MLLMs' ability to interpret and explain cinematographic techniques. For the generation task, we assess advanced video generation models on their capacity to reconstruct cinema-quality camera movements given conditions such as textual prompts or keyframes. We conduct a large-scale evaluation on 15+ MLLMs and 5+ video generation models. Our results offer insights into the limitations of current models and future directions for cinematography understanding and generation in automatically film production and appreciation. The code and benchmark can be accessed at https://github.com/PRIS-CV/CineTechBench.

📄 PDF Abstract BibTeX arXiv:2505.15145

Code (1)

pris-cv/cinetechbench 공식 구현 pytorch

Tasks

Video Generation

Similar Papers 제목 키워드 기반

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning

2026-06-23 · Xinyu Mao, Yuhui Zeng, Xiaokun Liu, Wenyu Qin 외 arxiv

Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot size, depth of field, composition, and shooting angle. This capability is importan…

Reinforcement LearningVideo CaptioningVideo Generation

Natural Language Camera Movement Understanding

2026-07-03 · Yuwen Tan, Joey Huang, Jin Huang, Haoxiang Li 외 arxiv

Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we demonstrate that existing vision-language models (VLMs) fail this t…

Video Generation

DramaDirector: Geometry-Guided Short Drama Generation

2026-06-23 · Hengji Zhou, Sijie Liu, Jianrun Chen, Xingchen Zou 외 arxiv

Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that prompt-level or text-only video generation pipelines struggle to meet. We study plo…

Video Generation

CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents

2026-07-29 · Qianru Li, Xuyang Chen, Erkin Türköz, Lu Liu 외 arxiv

Automatically generating cinematically expressive camera trajectories through 3D scenes from natural language descriptions is a challenging task of high practical value, with applications ranging from real-estate adverti…

Trajectory PlanningCollision AvoidanceSpatial Reasoning

ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions

2025-12-11 · Xiaoxue Wu, Xinyuan Chen, Yaohui Wang, Yu Qiao arxiv

Shot transitions play a pivotal role in multi-shot video generation, as they determine the overall narrative expression and the directorial design of visual storytelling. However, recent progress has primarily focused on…

Visual StorytellingVideo Generation