paper-with-me

홈 › Papers

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games

2026-05-05 · Connacher Murphy arxiv

Static capabilities benchmarks suffer from saturation and contamination, making it difficult to track capabilities progress over time. We introduce Agent Island, a multiplayer simulation environment in which language-model agents compete in a game of interagent cooperation, conflict, and persuasion. The environment yields a dynamic benchmark designed to mitigate both saturation and contamination; new models can always outperform the current leading player in this winner-take-all game, and agents compete against other adaptive agents rather than face a fixed task set. We rank players with a Bayesian Plackett-Luce model, allowing us to quantify uncertainty in player skill. In 999 games involving 49 unique models, openai/gpt-5.5 dominates its peers with a posterior mean skill of 5.64, compared with 3.10 for the second-ranked model, openai/gpt-5.2, and 2.86 for the third-ranked model, openai/gpt-5.3-codex. We release the game logs as a dataset for analyses of model behavior. As an example, we investigate same-provider preference in final-round votes and find that models are 8.3 p.p. more likely to support a same-provider finalist than finalists from other providers. This preference is not uniform across providers: among separately estimated providers, the effect is strongest for OpenAI models and weakest for Anthropic models.

📄 PDF Abstract BibTeX arXiv:2605.04312

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM Benchmark Datasets Should Be Contamination-Resistant

2026-05-19 · Ali Al-Lawati, Jason Lucas, Dongwon Lee, Suhang Wang arxiv

Benchmark datasets are critical for reproducible, reliable, and discriminative evaluation of LLMs. However, recent studies reveal that many benchmark datasets are included in pretraining corpora, i.e., $\textit{contamina…

Towards Contamination Resistant Benchmarks

2025-05-13 · Rahmatullah Musawi, Sheng Lu

The rapid development of large language models (LLMs) has transformed the landscape of natural language processing. Evaluating LLMs properly is crucial for understanding their potential and addressing concerns such as sa…

Toward Evaluation Frameworks for Multi-Agent Scientific AI Systems

2026-03-18 · Marcin Abram arxiv

We analyze the challenges of benchmarking scientific (multi)-agentic systems, including the difficulty of distinguishing reasoning from retrieval, the risks of data/model contamination, the lack of reliable ground truth …

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

2026-08-24 · Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan 외 arxiv

Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting ag…

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation

2026-05-29 · Andreas Haupt, Justin Hartenstein, Anka Reuel, Mykel Kochenderfer 외 arxiv

AI benchmarks have well-documented limitations, with prior work examining contamination, saturation, and construct underspecification. Aggregation has received far less attention: benchmarks are typically summarized by u…