paper-with-me

Papers

Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation

2025-09-30 · Agneet Chatterjee, Rahim Entezari, Maksym Zhuravinskyi, Maksim Lapin, Reshinth Adithyan, Amit Raj, Chitta Baral, Yezhou Yang, Varun Jampani arxiv

Recent advances in video generation have enabled high-fidelity video synthesis from user provided prompts. However, existing models and benchmarks fail to capture the complexity and requirements of professional video generation. Towards that goal, we introduce Stable Cinemetrics, a structured evaluation framework that formalizes filmmaking controls into four disentangled, hierarchical taxonomies: Setup, Event, Lighting, and Camera. Together, these taxonomies define 76 fine-grained control nodes grounded in industry practices. Using these taxonomies, we construct a benchmark of prompts aligned with professional use cases and develop an automated pipeline for prompt categorization and question generation, enabling independent evaluation of each control dimension. We conduct a large-scale human study spanning 10+ models and 20K videos, annotated by a pool of 80+ film professionals. Our analysis, both coarse and fine-grained reveal that even the strongest current models exhibit significant gaps, particularly in Events and Camera-related controls. To enable scalable evaluation, we train an automatic evaluator, a vision-language model aligned with expert annotations that outperforms existing zero-shot baselines. SCINE is the first approach to situate professional video generation within the landscape of video generative models, introducing taxonomies centered around cinematic controls and supporting them with structured evaluation pipelines and detailed analyses to guide future research.

📄 PDF Abstract BibTeX arXiv:2509.26555

Code (0)

등록된 구현이 없습니다.

Tasks

Question GenerationVideo Generation

Similar Papers 제목 키워드 기반

FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs

2025-10-10 · Yan Wang, Keyi Wang, Shanshan Yang, Jaisal Patel 외 arxiv

Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports are filed in XBRL, a structured XML form…

Information ExtractionMathematical Reasoning

BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing

2026-04-06 · Kaiwen Wang, Kaili Zheng, Rongrong Deng, Yiming Shi 외 arxiv

Recent multimodal large language models (MLLMs) have shown strong capabilities in general video understanding, driving growing interest in automatic sports commentary generation. However, existing benchmarks for this tas…

Responsible Evaluation of AI for Mental Health

2026-01-20 · Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt 외 arxiv

Although artificial intelligence (AI) shows growing promise for mental health care, current approaches to evaluating AI tools in this domain remain fragmented and poorly aligned with clinical practice, social context, an…

Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents

2026-08-06 · Tao Wang, Qihao Yang, Rongjiao Liang, Lianghong Lin 외 arxiv

Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T stan…

Question Answering

"It's a conversation, not a quiz": A Risk Taxonomy and Reflection Tool for LLM Adoption in Public Health

2024-11-04 · Jiawei Zhou, Amy Z. Chen, Darshi Shah, Laura Schwab Reese 외

Recent breakthroughs in large language models (LLMs) have generated both interest and concern about their potential adoption as accessible information sources or communication tools across different domains. In public he…