paper-with-me

Papers

Deprecating Benchmarks: Criteria and Framework

2025-07-08 · Ayrton San Joaquin, Rokas Gipiškis, Leon Staufer, Ariel Gil arxiv

As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance on when and how benchmarks should be deprecated once they cease to effectively perform their purpose. This risks benchmark scores over-valuing model capabilities, or worse, obscuring capabilities and safety-washing. Based on a review of benchmarking practices, we propose criteria to decide when to fully or partially deprecate benchmarks, and a framework for deprecating benchmarks. Our work aims to advance the state of benchmarking towards rigorous and quality evaluations, especially for frontier models, and our recommendations are aimed to benefit benchmark developers, benchmark users, AI governance actors (across governments, academia, and industry panels), and policy makers.

📄 PDF Abstract BibTeX arXiv:2507.06434

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An LSTM-Based Deep Learning Approach for Detecting Self-Deprecating Sarcasm in Textual Data

2019-12-01 · ICON 2019 12 · Ashraf Kamal, Muhammad Abulaish

Self-deprecating sarcasm is a special category of sarcasm, which is nowadays popular and useful for many real-life applications, such as brand endorsement, product campaign, digital marketing, and advertisement. The self…

Deep LearningMarketingSarcasm Detection

A Framework for Deprecating Datasets: Standardizing Documentation, Identification, and Communication

2021-10-18 · Alexandra Sasha Luccioni, Frances Corry, Hamsini Sridharan, Mike Ananny 외

Datasets are central to training machine learning (ML) models. The ML community has recently made significant improvements to data stewardship and documentation practices across the model development life cycle. However,…

Agent Benchmarks Fail Public Sector Requirements

2026-01-28 · Jonathan Rystrøm, Chris Schmitz, Karolina Korgul, Jan Batzner 외 arxiv

Deploying Large Language Model-based agents (LLM agents) in the public sector requires assuring that they meet the stringent legal, procedural, and structural requirements of public-sector institutions. Practitioners and…

Blending Pruning Criteria for Convolutional Neural Networks

2021-07-11 · wei he, Zhongzhan Huang, Mingfu Liang, Senwei Liang 외

The advancement of convolutional neural networks (CNNs) on various vision applications has attracted lots of attention. Yet the majority of CNNs are unable to satisfy the strict requirement for real-world deployment. To …

ClusteringNetwork Pruning

DEL: Digit Entropy Loss for Numerical Learning of Large Language Models

2026-05-19 · Zhaohui Zheng, Chenhang He, Shihao Wang, Yuxuan Li 외 arxiv

Number prediction stands as a fundamental capability of large language models (LLMs) in mathematical problem-solving and code generation. The widely adopted maximum likelihood estimation (MLE) for LLM training is not tai…

Mathematical ReasoningCode Generation