paper-with-me

홈 › Papers

HarnessEval-W: Agentifying the Evaluation of Visual Worlds

2026-08-17 · Weiliang Chen, Haowen Sun, Jun Gao, Jiawei Chi, Hanyang Wang, Qiyu Dai, Yihao Li, Hao Li, Jingnan Gao, Yi-Hsin Hung, Xingzhuo Guo, Shangchen Miao, Zhiyuan Shi, Xiang Li, Fengrui Tian, Weihua Du, Ziqi Huang, Shenyuan Gao, Siqiao Huang, Mingyu Liu, Yifei Li, Shizun Wang, Xi Wang, Tianqi Zhang, Xue Luo, Xiyin Ren, Jinshan Ren, Xiaoyang Shen, Xiaobo Hu, Zhiyang Dou, Mingyu Ding, Yichao Yan, Xinchao Wang, Yizhou Wang, Shilong Liu, Wenzhao Zheng, Yueqi Duan, Yuan Gong, Ziwei Liu, Ming-Yu Liu, Jialong Wu, Jiangran Lyu, Fangfu Liu arxiv

A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations naturally, yet no existing benchmark automates this capability: metrics are computed brute-force, leaving no reasoning chain that can be examined or verified. We introduce HarnessEval-W, an agentified evaluation pipeline that brings the harness paradigm from the LLM ecosystem to world model benchmarking. Rather than applying a fixed rubric, HarnessEval-W interprets the context of each evaluation case, decomposes the evaluation question into measurable subproblems, and spawns specialized sub-agents, each equipped with tailored context and diagnostic tools to reason over its own subproblem. The parent agent then validates the gathered evidence and summarizes it into the final verdict. This hierarchical workflow turns every evaluation into a transparent evidence tree whose complete reasoning chain justifies the result. We apply HarnessEval-W to 18 representative world models over 330 evaluation cases. Its judgments closely align with human preferences while providing verifiable, fine-grained diagnoses of every generated rollout. We open-source the full pipeline as a live benchmark and invite the broad community to contribute to grow new skills and evaluation cases as world models evolve.

📄 PDF Abstract BibTeX arXiv:2608.16859

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

WorldScribe: Towards Context-Aware Live Visual Descriptions

2024-08-13 · Ruei-Che Chang, Yuxuan Liu, Anhong Guo

Automated live visual descriptions can aid blind people in understanding their surroundings with autonomy and independence. However, providing descriptions that are rich, contextual, and just-in-time has been a long-stan…

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

2025-02-06 · Jack Hong, Shilin Yan, Jiayin Cai, XiaoLong Jiang 외

In this paper, we introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSens…

Video Understanding

WorldScore: A Unified Evaluation Benchmark for World Generation

2025-04-01 · Haoyi Duan, Hong-Xing Yu, Sirui Chen, Li Fei-Fei 외

We introduce the WorldScore benchmark, the first unified benchmark for world generation. We decompose world generation into a sequence of next-scene generation tasks with explicit camera trajectory-based layout specifica…

Scene GenerationVideo Generation

WorldSimBench: Towards Video Generation Models as World Simulators

2024-10-23 · Yiran Qin, Zhelun Shi, Jiwen Yu, Xijun Wang 외

Recent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to…

Autonomous DrivingRobot ManipulationVideo Generation

Code2Worlds: Empowering Coding LLMs for 4D World Generation

2026-02-12 · Yi Zhang, Yunshuang Wang, Zeyu Zhang, Hao Tang arxiv

Achieving spatial intelligence requires moving beyond visual plausibility to build world simulators grounded in physical laws. While coding LLMs have advanced static 3D scene generation, extending this paradigm to 4D dyn…

Scene GenerationCode Generation