paper-with-me

홈 › Papers

V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models

2025-11-20 · Yang Luo, Xuanlei Zhao, Baijiong Lin, Lingting Zhu, Liyao Tang, Yuqi Liu, Ying-Cong Chen, Shengju Qian, Xin Wang, Yang You arxiv

Recent progress in generative video models, such as Veo-3, has shown surprising zero-shot reasoning abilities, creating a growing need for systematic and reliable evaluation. We introduce V-ReasonBench, a benchmark designed to assess video reasoning across four key dimensions: structured problem-solving, spatial cognition, pattern-based inference, and physical dynamics. The benchmark is built from both synthetic and real-world image sequences and provides a diverse set of answer-verifiable tasks that are reproducible, scalable, and unambiguous. Evaluations of six state-of-the-art video models reveal clear dimension-wise differences, with strong variation in structured, spatial, pattern-based, and physical reasoning. We further compare video models with strong image models, analyze common hallucination behaviors, and study how video duration affects Chain-of-Frames reasoning. Overall, V-ReasonBench offers a unified and reproducible framework for measuring video reasoning and aims to support the development of models with more reliable, human-aligned reasoning skills.

📄 PDF Abstract BibTeX arXiv:2511.16668

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Similar Papers 제목 키워드 기반

VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?

2025-05-29 · Yuanxin Liu, Kun Ouyang, HaoNing Wu, Yi Liu 외

Recent studies have shown that long chain-of-thought (CoT) reasoning can significantly enhance the performance of large language models (LLMs) on complex tasks. However, this benefit is yet to be demonstrated in the doma…

Video Understanding

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

2026-05-11 · Keming Wu, Yijing Cui, Wenhan Xue, Qijie Wang 외 arxiv

Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving into "world simulators." Yet the community still lacks a benchmark th…

Video Generation

T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation

2025-08-24 · Kaiyue Sun, Rongyao Fang, Chengqi Duan, Xian Liu 외 arxiv

We propose T2I-ReasonBench, a benchmark evaluating reasoning capabilities of text-to-image (T2I) models. It consists of four dimensions: Idiom Interpretation, Textual Image Design, Entity-Reasoning and Scientific-Reasoni…

Text-to-Image Generation

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning

2025-12-08 · Nearchos Potamitis, Vansh Ramani, Har Ashish Arora, Dhairya Kuchhal 외 arxiv

Benchmark scores for LLM reasoning systems are reported as single numbers, yet the same model, strategy, and task can produce meaningfully different answers and costs across repeated executions, even under greedy decodin…

CXReasonBench: A Benchmark for Evaluating Structured Diagnostic Reasoning in Chest X-rays

2025-05-23 · Hyungyung Lee, Geon Choi, Jung-Oh Lee, Hangyul Yoon 외

Recent progress in Large Vision-Language Models (LVLMs) has enabled promising applications in medical tasks, such as report generation and visual question answering. However, existing benchmarks focus mainly on the final…

DiagnosticQuestion AnsweringVisual GroundingVisual Question Answering