paper-with-me

홈 › Papers

Train-before-Test Harmonizes Language Model Rankings

2025-07-07 · Guanhua Zhang, Ricardo Dominguez-Olmedo, Moritz Hardt arxiv

Existing language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills. This dilemma of conflicting rankings hampers model selection, clouds model comparisons, and adds confusion to a growing ecosystem of competing models. In this paper, we take a different perspective on model comparison: instead of relying on out-of-the-box performance via direct evaluation, we compare model potential by providing each model with identical benchmark-specific fine-tuning before evaluation. We call this approach train-before-test. Our primary contribution is a comprehensive empirical evaluation of model potential across 24 benchmarks and 61 models. First, we demonstrate that model potential rankings obtained through train-before-test exhibit remarkable consistency across all benchmarks. Whereas traditional rankings demonstrate little external validity under direct evaluation, they enjoy a significant degree of external validity when applying train-before-test: model potential rankings transfer gracefully from one benchmark to another. Second, train-before-test restores the connection between perplexity and downstream task performance, lost under direct evaluation. Remarkably, even pre-finetuning perplexity of a base model predicts post-finetuning downstream performance, suggesting that ranking consistency reflects inherent model potential rather than fine-tuning artifacts. Finally, train-before-test reduces the model-score matrix to essentially rank one, indicating that model potential is dominated by one latent factor, uncovered by train-before-test. Our work supports the recommendation to make train-before-test a default component of LLM benchmarking.

📄 PDF Abstract BibTeX arXiv:2507.05195

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Self-critical Sequence Training for Automatic Speech Recognition

2022-04-13 · Chen Chen, Yuchen Hu, Nana Hou, Xiaofeng Qi 외

Although automatic speech recognition (ASR) task has gained remarkable success by sequence-to-sequence models, there are two main mismatches between its training and testing that might lead to performance degradation: 1)…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Reinforcement Learning (RL)speech-recognition+1

Associative Measures and Multi-word Unit Extraction in Turkish

2015-07-15 · Umit Mersinli

Associative measures are "mathematical formulas determining the strength of association between two or more words based on their occurrences and cooccurrences in a text corpus" (Pecina, 2010, p. 138). The purpose of this…

Sentence

Intertemporal Connections Between Query Suggestions and Search Engine Results for Politics Related Queries

2018-12-20 · Malte Bonart, Philipp Schaer

This short paper deals with the combination and comparison of two data sources: Search engine results and query suggestions for 16 terms related to political candidates and parties. The data was collected before the fede…

Partial Rankings of Optimizers

2024-02-26 · Julian Rodemann, Hannah Blocher

We introduce a framework for benchmarking optimizers according to multiple criteria over various test functions. Based on a recently introduced union-free generic depth function for partial orders/rankings, it fully expl…

Benchmarking

Crowdsourcing Relative Rankings of Multi-Word Expressions: Experts versus Non-Experts

2022-06-17 · David Alfter, Therese Lindström Tiedemann, Elena Volodina

In this study we investigate to which degree experts and non-experts agree on questions of difficulty in a crowdsourcing experiment. We ask non-experts (second language learners of Swedish) and two groups of experts (tea…