paper-with-me

Papers

World Consistency Score: A Unified Metric for Video Generation Quality

2025-07-31 · Akshat Rakheja, Aarsh Ashdhir, Aryan Bhattacharjee, Vanshika Sharma arxiv

We introduce World Consistency Score (WCS), a novel unified evaluation metric for generative video models that emphasizes internal world consistency of the generated videos. WCS integrates four interpretable sub-components - object permanence, relation stability, causal compliance, and flicker penalty - each measuring a distinct aspect of temporal and physical coherence in a video. These submetrics are combined via a learned weighted formula to produce a single consistency score that aligns with human judgments. We detail the motivation for WCS in the context of existing video evaluation metrics, formalize each submetric and how it is computed with open-source tools (trackers, action recognizers, CLIP embeddings, optical flow), and describe how the weights of the WCS combination are trained using human preference data. We also outline an experimental validation blueprint: using benchmarks like VBench-2.0, EvalCrafter, and LOVE to test WCS's correlation with human evaluations, performing sensitivity analyses, and comparing WCS against established metrics (FVD, CLIPScore, VBench, FVMD). The proposed WCS offers a comprehensive and interpretable framework for evaluating video generation models on their ability to maintain a coherent "world" over time, addressing gaps left by prior metrics focused only on visual fidelity or prompt alignment.

📄 PDF Abstract BibTeX arXiv:2508.00144

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

WorldMark: A Unified Benchmark Suite for Interactive Video World Models

2026-04-23 · Xiaojie Xu, Zhengyuan Lin, Kang He, Yukang Feng 외 arxiv

Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated on its own benchmark with private scenes and trajectories, making fair cross-model co…

Video Generation

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing

2025-01-13 · Varun Biyyala, Bharat Chanderprakash Kathuria, Jialu Li, Youshan Zhang

Video editing models have advanced significantly, but evaluating their performance remains challenging. Traditional metrics, such as CLIP text and image scores, often fall short: text scores are limited by inadequate tra…

Objectobject-detectionObject DetectionObject Tracking+1

FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction

2025-09-25 · Yixiang Dai, Fan Jiang, Chiyu Wang, Mu Xu 외 arxiv

High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established stron…

Novel View SynthesisVideo Generation

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning

2026-07-20 · Ting Huang, Zhenyu Zhang, Wenyuan Huang, Jian Yang 외 hf

Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multi…

Video Question AnsweringReinforcement LearningSpatial Reasoning

The Trinity of Consistency as a Defining Principle for General World Models

2026-02-26 · Jingxuan Wei, Siyuan Li, Yuhang Xu, Zheng Sun 외 arxiv

The construction of World Models capable of learning, simulating, and reasoning about objective physical laws constitutes a foundational challenge in the pursuit of Artificial General Intelligence. Recent advancements re…

Video Generation