paper-with-me

Papers

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

2026-05-28 · You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu, Zhixiang Wang arxiv

As video diffusion models (VDMs) advance toward world models, a key question arises: do they truly understand causality, or merely overfit to statistical temporal patterns? Existing benchmarks mostly rely on synthetic data, limiting real-world generalization due to the sim-to-real gap. We present YoCausal, a two-level benchmark inspired by the Violation of Expectation (VoE) paradigm from cognitive science. By temporally reversing real-world videos at zero cost as natural counterfactual samples, YoCausal establishes an arbitrarily extensible evaluation protocol. Level 1 introduces the Reverse Surprise Index (RSI), quantifying arrow-of-time perception via denoising loss. Level 2 introduces the Causality Cognition Index (CCI), which leverages a VLM to stratify datasets into causal and non-causal subsets, disentangling genuine causal reasoning from temporal bias. Evaluation of 13 state-of-the-art VDMs reveals that perceiving the arrow of time does not imply understanding causality, and a significant gap persists relative to human-level causal cognition.

📄 PDF Abstract BibTeX arXiv:2605.30346

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?

2025-10-09 · Zeqing Wang, Xinyu Wei, Bairui Li, Zhen Guo 외 arxiv

The recent rapid advancement of Text-to-Video (T2V) generation technologies are engaging the trained models with more world model ability, making the existing benchmarks increasingly insufficient to evaluate state-of-the…

Reasoning before Responding: Integrating Commonsense-based Causality Explanation for Empathetic Response Generation

2023-07-28 · Yahui Fu, Koji Inoue, Chenhui Chu, Tatsuya Kawahara

Recent approaches to empathetic response generation try to incorporate commonsense knowledge or reasoning about the causes of emotions to better understand the user's experiences and feelings. However, these approaches m…

Empathetic Response GenerationIn-Context LearningResponse Generation

MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning

2024-09-26 · Tieyuan Chen, Huabin Liu, Tianyao He, Yihang Chen 외

Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily executed in a question-answering paradi…

Causal DiscoveryCausal Discovery in Video Reasoning

MECD+: Unlocking Event-Level Causal Graph Discovery for Video Reasoning

2025-01-13 · Tieyuan Chen, Huabin Liu, Yi Wang, Yihang Chen 외

Video causal reasoning aims to achieve a high-level understanding of videos from a causal perspective. However, it exhibits limitations in its scope, primarily executed in a question-answering paradigm and focusing on br…

Causal DiscoveryCausal InferencecounterfactualCounterfactual Inference+3

T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation

2025-07-24 · Yubin Chen, Xuyang Guo, Zhenmei Shi, Zhao Song 외 arxiv

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains lar…

Text-to-Video Generation