paper-with-me

홈 › Papers

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

2026-06-30 · Ting-Bing Xu, Jiacheng Sui, Zhe Gao, Kewei Shi, Wenjin Yang, Zhicheng Liu, Zhaoxu Sun, Mingchao Sun, Hongyu Pan, Fan Jiang, Mu Xu, Qi Fan, Yang Gao, Yong Li, Baoquan Chen arxiv

Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics. We introduce WorldRoamBench, an open-world benchmark for long-horizon stability across four dimensions, each with tailored innovations: (i) Action: per-frame action metric bypassing cross-model semantic scale disparity and exposing failures hidden by trajectory; (ii) Vision: segment-based drift metric capturing non-monotonic mid-sequence collapse missed by start-vs-end comparisons; (iii) Physics: controllability-gated evaluation over mechanics, optics, and 3D consistency, scoring plausibility under faithful action execution; (iv) Memory: action-decoupled protocol evaluating scene memory via transition-localized 3D point-cloud reconstruction and subject memory via tracking-plus-VLM reasoning. The benchmark comprises 600+ test cases across Nature, Urban, and Indoor scenes in first/third-person views with WASD 10-60s continuous interaction. Evaluating 10+ open/closed-source models reveals none reliably satisfies all dimensions; even the best achieves only moderate scores. Advances on WorldRoamBench are steps toward IWMs that are stable, physically grounded, memory-faithful, and deployable in real-world applications.

📄 PDF Abstract BibTeX arXiv:2606.31672

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

2026-07-21 · Fan Jiang, Zhaoxu Sun, Mengchao Wang, Ziyu Zhu 외 hf

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet v…

Open-World Video Segmentation

2026-06-14 · Qing Su, Kaiyang Li, Yuan Zhuang, Fei Miao 외 arxiv

While video segmentation has advanced rapidly on short clips and closed-set benchmarks, open-world video segmentation remains largely unexplored. The challenge is twofold: (1) existing methods are not designed to support…

Open-World Video Segmentation

SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

2025-12-20 · Tue Le, Minh V. T. Thai, Dung Nguyen Manh, Huy Phan Nhat 외 arxiv

Existing benchmarks for AI coding agents focus on isolated, single-issue tasks such as fixing a bug or adding a small feature. However, real-world software engineering is a long-horizon endeavor: developers interpret hig…

NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama

2026-06-16 · Logan Mann, Abdur Rahman, Mohammad Saifullah, Taaha Kazi 외 arxiv

Long-form serialized audio drama, with arcs that run for 200 to 800 episodes, is a major creative medium and a setting where frontier large language models (LLMs) fail. We benchmark 21 models, spanning classical, fine-tu…

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

2026-06-05 · Siyi Chen, Hugo Hadfield, Alex Zook, Mikaela Angelina Uy 외 arxiv

Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning, executing, monitoring, and recovering from failures. We address th…