paper-with-me

Papers

ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

2025-06-26 · Hongbo Liu, Jingwen He, Yi Jin, Dian Zheng, Yuhao Dong, Fan Zhang, Ziqi Huang, Yinan He, Yangguang Li, WeiChao Chen, Yu Qiao, Wanli Ouyang, Shengjie Zhao, Ziwei Liu

Cinematography, the fundamental visual language of film, is essential for conveying narrative, emotion, and aesthetic quality. While recent Vision-Language Models (VLMs) demonstrate strong general visual understanding, their proficiency in comprehending the nuanced cinematic grammar embedded within individual shots remains largely unexplored and lacks robust evaluation. This critical gap limits both fine-grained visual comprehension and the precision of AI-assisted video generation. To address this, we introduce \textbf{ShotBench}, a comprehensive benchmark specifically designed for cinematic language understanding. It features over 3.5k expert-annotated QA pairs from images and video clips, meticulously curated from over 200 acclaimed (predominantly Oscar-nominated) films and spanning eight key cinematography dimensions. Our evaluation of 24 leading VLMs on ShotBench reveals their substantial limitations: even the top-performing model achieves less than 60\% average accuracy, particularly struggling with fine-grained visual cues and complex spatial reasoning. To catalyze advancement in this domain, we construct \textbf{ShotQA}, a large-scale multimodal dataset comprising approximately 70k cinematic QA pairs. Leveraging ShotQA, we develop \textbf{ShotVL} through supervised fine-tuning and Group Relative Policy Optimization. ShotVL significantly outperforms all existing open-source and proprietary models on ShotBench, establishing new \textbf{state-of-the-art} performance. We open-source our models, data, and code to foster rapid progress in this crucial area of AI-driven cinematic understanding and generation.

📄 PDF Abstract BibTeX arXiv:2506.21356

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningVideo Generation

Similar Papers 제목 키워드 기반

RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation

2025-10-02 · Hang Wu, Yujun Cai, Haonan Ge, Hongkai Chen 외 arxiv

Cinematography understanding refers to the ability to recognize not only the visual content of a scene but also the cinematic techniques that shape narrative meaning. This capability is attracting increasing attention, a…

A Benchmark for Omni-Modal Reasoning in Long Videos

2025-12-18 · Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Jinxing Zhou, Sahal Shaji Mullappilly 외 arxiv

Long-form omni-modal video understanding requires integrating vision, speech, and ambient audio with coherent long-context reasoning. Existing video benchmarks often trade off temporal scale, modality coverage, open-ende…

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

2026-05-22 · Songlin Yang, Haobin Zhong, Ruilin Zhang, Xiaotong Zhao 외 arxiv

The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community transitions towards Reinforcement Learning…

Reinforcement LearningVideo Generation

Fine-Tuning Open Video Generators for Cinematic Scene Synthesis: A Small-Data Pipeline with LoRA and Wan2.1 I2V

2025-10-31 · Meftun Akarsu, Kerem Catay, Sedat Bin Vedat, Enes Kutay Yarkan 외 arxiv

We present a practical pipeline for fine-tuning open-source video diffusion transformers to synthesize cinematic scenes for television and film production from small datasets. The proposed two-stage process decouples vis…

Seeking Universal Shot Language Understanding Solutions

2026-03-19 · Haoxin Liu, Harshavardhan Kamarthi, Zhiyuan Zhao, Hongjie Chen 외 arxiv

Shot language understanding (SLU) is crucial for cinematic analysis but remains challenging due to its diverse cinematographic dimensions and subjective expert judgment. While vision-language models (VLMs) have shown str…