paper-with-me

홈 › Papers

Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking

2026-05-07 · Yang Xu, Jiefu Zhang, Haixiang Sun, Zihan Zhou, Tianyu Cao, Vaneet Aggarwal arxiv

Adaptive prompt and program search makes LLM evaluation selection-sensitive. Once benchmark items are reused inside tuning, the observed winner's score need not estimate the fresh-data performance of the full tune-then-deploy procedure. We study inference for this procedure-level target under explicit tuning budgets. We propose SIREN, a selection-aware repeated-split reporting protocol that freezes the post-search shortlist, separates splitwise selection from held-out evaluation, and uses an item-level Gaussian multiplier bootstrap for uncertainty quantification. In a fixed-shortlist regime with smooth stabilized selection, the estimator admits a first-order item-level representation, and the bootstrap yields valid simultaneous inference on a finite budget grid. This supports confidence intervals for procedure-performance curves and pre-specified equal-budget and cross-budget comparisons. Controlled simulations and MMLU-Pro tuning experiments show that winner-based reporting can be optimistic and can change deployment conclusions, while SIREN remains close to the finite-sample reporting target.

📄 PDF Abstract BibTeX arXiv:2605.05973

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Inference on Multiple Winners with Applications to Microcredit and Economic Mobility

2024-10-24 · Andreas Petrou-Zeniou, Azeem M. Shaikh

While policymakers and researchers are often concerned with conducting inference based on a data-dependent selection, a strictly larger class of inference problems arises when considering multiple data-dependent selectio…

Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching

2026-02-09 · Hamsa Bastani, Osbert Bastani, Bryce McLaughlin arxiv

A major challenge in data-driven decision-making is accurate policy evaluation-i.e., guaranteeing that a learned decision-making policy achieves the promised benefits. A popular strategy is model-based policy evaluation,…

A Flexible Defense Against the Winner's Curse

2024-11-27 · Tijana Zrnic, William Fithian

Across science and policy, decision-makers often need to draw conclusions about the best candidate among competing alternatives. For instance, researchers may seek to infer the effectiveness of the most successful treatm…

Selection biasvalid

Cursed yet Satisfied Agents

2021-04-02 · YiLing Chen, Alon Eden, Juntao Wang

In real life auctions, a widely observed phenomenon is the winner's curse -- the winner's high bid implies that the winner often over-estimates the value of the good for sale, resulting in an incurred negative utility. T…

Beating the Winner's Curse via Inference-Aware Policy Optimization

2025-10-20 · Hamsa Bastani, Osbert Bastani, Bryce McLaughlin arxiv

There has been a surge of recent interest in automatically learning policies to target treatment decisions based on rich individual covariates. In addition, practitioners want confidence that the learned policy has bette…