paper-with-me

Papers

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation

2025-05-01 · D. Sculley, Will Cukierski, Phil Culliton, Sohier Dane, Maggie Demkin, Ryan Holbrook, Addison Howard, Paul Mooney, Walter Reade, Megan Risdal, Nate Keating

In this position paper, we observe that empirical evaluation in Generative AI is at a crisis point since traditional ML evaluation and benchmarking strategies are insufficient to meet the needs of evaluating modern GenAI models and systems. There are many reasons for this, including the fact that these models typically have nearly unbounded input and output spaces, typically do not have a well defined ground truth target, and typically exhibit strong feedback loops and prediction dependence based on context of previous model outputs. On top of these critical issues, we argue that the problems of leakage and contamination are in fact the most important and difficult issues to address for GenAI evaluations. Interestingly, the field of AI Competitions has developed effective measures and practices to combat leakage for the purpose of counteracting cheating by bad actors within a competition setting. This makes AI Competitions an especially valuable (but underutilized) resource. Now is time for the field to view AI Competitions as the gold standard for empirical rigor in GenAI evaluation, and to harness and harvest their results with according value.

📄 PDF Abstract BibTeX arXiv:2505.00612

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingPosition

Similar Papers 제목 키워드 기반

On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert Ratings

2021-09-07 · EMNLP 2021 11 · Peter Jansen, Kelly Smith, Dan Moreno, Huitzilin Ortiz

Building compositional explanations requires models to combine two or more facts that, together, describe why the answer to a question is correct. Typically, these "multi-hop" explanations are evaluated relative to one (…

valid

Velocity and stroke rate reconstruction of canoe sprint team boats based on panned and zoomed video recordings

2026-02-26 · Julian Ziegler, Daniel Matthes, Finn Gerdts, Patrick Frenzel 외 arxiv

Pacing strategies, defined by velocity and stroke rate profiles, are essential for peak performance in canoe sprint. While GPS is the gold standard for analysis, its limited availability necessitates automated video-base…

GhoSt-PV: A Representative Gold Standard of German Particle Verbs

2016-12-01 · WS 2016 12 · Stefan Bott, Nana Khvtisavrishvili, Max Kisselew, Sabine Schulte im Walde

German particle verbs represent a frequent type of multi-word-expression that forms a highly productive paradigm in the lexicon. Similarly to other multi-word expressions, particle verbs exhibit various levels of composi…

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

2026-09-02 · Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi, Somshubra Majumdar 외 hf

Competitive programming has become a key test of large language model reasoning, with international competitions such as IOI and ICPC representing its most challenging settings. We present an end-to-end specialization pi…

Reinforcement Learning

Universal NER: A Gold-Standard Multilingual Named Entity Recognition Benchmark

2023-11-15 · arXiv 2023 11 · Stephen Mayhew, Terra Blevins, Shuheng Liu, Marek Šuppa 외

We introduce Universal NER (UNER), an open, community-driven project to develop gold-standard NER benchmarks in many languages. The overarching goal of UNER is to provide high-quality, cross-lingually consistent annotati…

Cross-Lingual NERMultilingual Named Entity Recognitionnamed-entity-recognitionNamed Entity Recognition+2