paper-with-me

홈 › Papers

WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation

2026-05-25 · Kaining Ying, Hengrui Hu, Siyu Ren, Jiamu Li, Fengjiao Chen, Ziwen Wang, Xuezhi Cao, Xunliang Cai, Henghui Ding arxiv

Interactive world models are advancing rapidly, yet existing benchmarks cover only part of the required competencies, leaving no unified standard for systematic evaluation. To fill this gap, we introduce WBench, a comprehensive multi-turn benchmark for interactive world model evaluation along five dimensions, namely video quality, setting adherence, interaction adherence, consistency, and physics compliance. WBench contains 289 test cases and 1,058 interaction turns, where each case specifies a world setting and a multi-turn interaction sequence, covering diverse scenes, styles, subjects, and both first- and third-person perspectives, together with four interaction types, including navigation, subject action, event editing, and perspective switching. For navigation, WBench unifies text, 6-DoF pose, and discrete-action control, enabling evaluation of models with different native input interfaces. Evaluation uses 22 automatic sub-metrics that combine specialist vision models with large multimodal models, and all metrics are validated against human judgments. Across 20 state-of-the-art models, we find that no single model performs strongly across all dimensions. We provide detailed diagnostic insights into the characteristic strengths, weaknesses, and open challenges of each model. Code and data are available at https://github.com/meituan-longcat/WBench.

📄 PDF Abstract BibTeX arXiv:2605.25874

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation

2025-04-30 · Sizhe Wang, Zhengren Wang, Dongsheng Ma, Yongan Yu 외

Modern software development demands code that is maintainable, testable, and scalable by organizing the implementation into modular components with iterative reuse of existing codes. We formalize this iterative, multi-tu…

Code Generation

DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation

2026-06-11 · Li Zhang, Yuzhen Shi, Yiran Hu, Jingwen Zhang 외 arxiv

Lawyer-client consultation is a critical starting point for legal services. Effective legal assistance hinges on eliciting sufficient and truthful information from clients in order to devise strategies that best protect …

Legal Reasoning

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

2026-07-09 · Zhekai Chen, Chengqi Duan, Kaiyue Sun, Bohao Li 외 arxiv

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. Ho…

StructFlowBench: A Structured Flow Benchmark for Multi-turn Instruction Following

2025-02-20 · Jinnan Li, Jinzhe Li, Yue Wang, Yi Chang 외

Multi-turn instruction following capability constitutes a core competency of large language models (LLMs) in real-world applications. Existing evaluation benchmarks predominantly focus on fine-grained constraint satisfac…

Instruction Following

FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

2023-10-31 · Yuxin Jiang, YuFei Wang, Xingshan Zeng, Wanjun Zhong 외

The ability to follow instructions is crucial for Large Language Models (LLMs) to handle various real-world applications. Existing benchmarks primarily focus on evaluating pure response quality, rather than assessing whe…

Instruction Following