paper-with-me

홈 › Papers

Leaderboard Incentives: Model Rankings under Strategic Post-Training

2026-03-09 · Yatong Chen, Guanhua Zhang, Moritz Hardt arxiv

Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmaxxing or training on the test task. In this work, we initiate a principled study of the incentive structure that benchmarks induce. We model benchmarking as a Stackelberg game between a benchmark designer who chooses an evaluation protocol and multiple model developers who compete simultaneously in a subgame given by the designer's choice. Each competitor has a model of unknown latent quality and can inflate its observed score by allocating resources to benchmark-specific improvements. First, we prove that current benchmarks induce games for which no Nash equilibrium between model developers exists. This result suggests one explanation for why current practice leads to misaligned incentives, prompting model developers to strategize in opaque ways. However, we prove that under mild conditions, a recently proposed evaluation protocol, called tune-before-test, induces a benchmark with a unique Nash equilibrium that ranks models by latent quality. This positive result demonstrates that benchmarks need not set bad incentives, even if current evaluations do.

📄 PDF Abstract BibTeX arXiv:2603.08371

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

How to Host a Data Competition: Statistical Advice for Design and Analysis of a Data Competition

2019-01-16 · Christine M. Anderson-Cook, Kary L. Myers, Lu Lu, Michael L. Fugate 외

Data competitions rely on real-time leaderboards to rank competitor entries and stimulate algorithm improvement. While such competitions have become quite popular and prevalent, particularly in supervised learning format…

LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts

2025-09-20 · Junhao Chen, Jingbo Sun, Xiang Li, Haidong Xin 외 arxiv

As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important. To fully assess LLM intelligence, it is crucial to examine their i…

Evaluating Strategic Reasoning in Forecasting Agents

2026-04-28 · Tom Liptay, Dan Schwarz, Rafael Poyiadzi, Jack Wildman 외 arxiv

Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to the Future 2 (BTF-2), 1,417 pastcasting questions with a frozen 15M-d…

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

2026-05-24 · Michael Hardy, Anka Reuel, Lijin Zhang, Jodi M. Casabianca 외 arxiv

While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when rankings reflect genuine capability differences…

Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard

2024-07-10 · Oguzhan Topsakal, Colby Jacob Edell, Jackson Bailey Harper

We introduce a novel and extensible benchmark for large language models (LLMs) through grid-based games such as Tic-Tac-Toe, Connect Four, and Gomoku. The open-source game simulation code, available on GitHub, allows LLM…