paper-with-me

홈 › Papers

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

2026-08-04 · Ine Gevers, Walter Daelemans arxiv

Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability of widely adopted commonsense benchmarks, we evaluate 23 models from six families on four established commonsense benchmarks, four reworked variants, three non-commonsense controls, and eight downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning. We compare model rankings, compute controlled correlations, and use leave-one-family-out cross-validation to assess the criterion validity of commonsense benchmarks. Our results show that revised benchmarks largely preserve original model rankings and do not improve downstream predictive power. Commonsense benchmarks show consistent cross-family predictive validity for only a narrow subset of downstream tasks, with smaller or metric-specific gains elsewhere. Overall, standardized commonsense benchmarks provide task-dependent rather than broad evidence of downstream commonsense competence.

📄 PDF Abstract BibTeX arXiv:2608.03340

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Fluid Language Model Benchmarking

2025-09-14 · Valentin Hofmann, David Heineman, Ian Magnusson, Kyle Lo 외 arxiv

Language model (LM) benchmarking faces several challenges: comprehensive evaluations are costly, benchmarks often fail to measure the intended capabilities, and evaluation quality can degrade due to labeling errors and b…

Predictive Models from Quantum Computer Benchmarks

2023-05-15 · Daniel Hothem, Jordan Hines, Karthik Nataraj, Robin Blume-Kohout 외

Holistic benchmarks for quantum computers are essential for testing and summarizing the performance of quantum hardware. However, holistic benchmarks -- such as algorithmic or randomized benchmarks -- typically do not pr…

Benchmarkingimage-classificationImage ClassificationTransfer Learning

The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models

2025-10-27 · Timo Freiesleben, Sebastian Zezulka arxiv

Predictive benchmarking, the evaluation of machine learning models based on predictive performance and competitive ranking, is a central epistemic practice in machine learning research and an increasingly prominent metho…

Image Classification

$\texttt{ACCORD}$: Closing the Commonsense Measurability Gap

2024-06-04 · François Roewer-Després, Jinyue Feng, Zining Zhu, Frank Rudzicz

We present $\texttt{ACCORD}$, a framework and benchmark suite for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs) through controlled, multi-hop counterfactuals. $\texttt{AC…

BenchmarkingCommon Sense ReasoningCounterfactual ReasoningLarge Language Model+1

CRoW: Benchmarking Commonsense Reasoning in Real-World Tasks

2023-10-23 · Mete Ismayilzada, Debjit Paul, Syrielle Montariol, Mor Geva 외

Recent efforts in natural language processing (NLP) commonsense reasoning research have yielded a considerable number of new datasets and benchmarks. However, most of these datasets formulate commonsense reasoning challe…

Benchmarking