paper-with-me

홈 › Papers

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

2026-08-13 · Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao arxiv

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.

📄 PDF Abstract BibTeX arXiv:2608.13552

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PlayWorld: Learning Robot World Models from Autonomous Play

2026-03-09 · Tenny Yin, Zhiting Mei, Zhonghe Zheng, Miyu Yamane 외 arxiv

Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot datasets, current state-of-the-art video…

Reinforcement Learning

A Coach-Player Framework for Dynamic Team Composition

2021-01-01 · Bo Liu, Qiang Liu, Peter Stone, Animesh Garg 외

In real-world multi-agent teams, agents with different capabilities may join or leave "on the fly" without altering the team's overarching goals. Coordinating teams with such dynamic composition remains a challenging pro…

Zero-shot Generalization

Sketchtopia: A Dataset and Foundational Agents for Benchmarking Asynchronous Multimodal Communication with Iconic Feedback

2025-01-01 · CVPR 2025 1 · Mohd Hozaifa Khan, Ravi Kiran Sarvadevabhatla

We introduce Sketchtopia, a large-scale dataset and AI framework designed to explore goal-driven, multimodal communication through asynchronous interactions in a Pictionary-inspired setup. Sketchtopia captures natura…

Benchmarking

Reflective Oracles: A Foundation for Classical Game Theory

2015-08-17 · Benja Fallenstein, Jessica Taylor, Paul F. Christiano

Classical game theory treats players as special---a description of a game contains a full, explicit enumeration of all players---even though in the real world, "players" are no more fundamentally special than rocks or cl…

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

2026-05-27 · Fangfu Liu, Kai He, Tianchang Shen, Tianshi Cao 외 arxiv

World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many generated environments require multi-agent …

Video Generation