paper-with-me

Papers

AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

2026-07-20 · Marjan Moodi, Xuankang Zhu, Fernando De Mesentier Silva, Harold Chaput, Mohammad Reza Taesiri hf

World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement direction is not specified in advance, unlike the engineering-to-spec tasks that dominate current agent benchmarks. We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided world-model starter under a fixed compute budget. The benchmark spans eight game environments under a unified structured-state representation--ground-truth entity state extracted from each game and consumed through a shared tensor format--which isolates dynamics modeling from perception and enables minutes-per-run iteration. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improve their starter on 63; in 91% of sessions the winning edit is a non-trivial research-style modification--a new objective, representation, rollout procedure, or architectural change--rather than a hyperparameter tweak. Our benchmark offers a setting in which frontier coding agents can be evaluated on open-ended research rather than engineering-to-spec problems.

📄 PDF Abstract BibTeX arXiv:2608.11216

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HumanVBench: Exploring Human-Centric Video Understanding Capabilities of MLLMs with Synthetic Benchmark Data

2024-12-23 · Ting Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding 외

In the domain of Multimodal Large Language Models (MLLMs), achieving human-centric video understanding remains a formidable challenge. Existing benchmarks primarily emphasize object and action recognition, often neglecti…

Action RecognitionVideo Understanding

Proof2Hybrid: Automatic Mathematical Benchmark Synthesis for Proof-Centric Problems

2025-08-04 · Yebo Peng, Zixiang Liu, Yaoming Li, Zhizhuo Yang 외 arxiv

Evaluating the mathematical capability of Large Language Models (LLMs) is a critical yet challenging frontier. Existing benchmarks fall short, particularly for proof-centric problems, as manual creation is unscalable and…

Open3DVQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space

2025-03-14 · Weichen Zhang, Zile Zhou, Zhiheng Zheng, Chen Gao 외

Spatial reasoning is a fundamental capability of embodied agents and has garnered widespread attention in the field of multimodal large language models (MLLMs). In this work, we propose a novel benchmark, Open3DVQA, to c…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+2

Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models

2025-09-30 · Yuansen Liu, Haiming Tang, Jinlong Peng, Jiangning Zhang 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the abs…

Scene Understanding

ThermoHands: A Benchmark for 3D Hand Pose Estimation from Egocentric Thermal Images

2024-03-14 · Fangqiang Ding, Yunzhou Zhu, Xiangyu Wen, Gaowen Liu 외

Designing egocentric 3D hand pose estimation systems that can perform reliably in complex, real-world scenarios is crucial for downstream applications. Previous approaches using RGB or NIR imagery struggle in challenging…

3D Hand Pose EstimationHand Pose EstimationPose Estimation