paper-with-me

홈 › Papers

Graders should cheat: privileged information enables expert-level automated evaluations

2025-02-16 · Jin Peng Zhou, Sébastien M. R. Arnold, Nan Ding, Kilian Q. Weinberger, Nan Hua, Fei Sha

Auto-evaluating language models (LMs), i.e., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and the cost associated with it. But this presents a paradox: how can we trust the grader LM, which is presumably weaker than the candidate LM, to assess problems that are beyond the frontier of the capabilities of either model or both? For instance, today's LMs struggle on graduate-level physics and Olympiad-level math, making them unreliable graders in these domains. We show that providing privileged information -- such as ground-truth solutions or problem-specific guidelines -- improves automated evaluations on such frontier problems. This approach offers two key advantages. First, it expands the range of problems where LMs graders apply. Specifically, weaker models can now rate the predictions of stronger models. Second, privileged information can be used to devise easier variations of challenging problems which improves the separability of different LMs on tasks where their performance is generally low. With this approach, general-purpose LM graders match the state of the art performance on RewardBench, surpassing almost all the specially-tuned models. LM graders also outperform individual human raters on Vibe-Eval, and approach human expert graders on Olympiad-level math problems.

📄 PDF Abstract BibTeX arXiv:2502.10961

Code (0)

등록된 구현이 없습니다.

Tasks

Math

Similar Papers 제목 키워드 기반

Learning by Cheating

2019-12-27 · Dian Chen, Brady Zhou, Vladlen Koltun, Philipp Krähenbühl

Vision-based urban driving is hard. The autonomous system needs to learn to perceive the world and act in it. We show that this challenging learning problem can be simplified by decomposing it into two stages. We first t…

Autonomous Driving

Cheating Automatic Short Answer Grading: On the Adversarial Usage of Adjectives and Adverbs

2022-01-20 · Anna Filighera, Sebastian Ochs, Tim Steuer, Thomas Tregel

Automatic grading models are valued for the time and effort saved during the instruction of large student bodies. Especially with the increasing digitization of education and interest in large-scale standardized testing,…

Adversarial Attackautomatic short answer gradingvalid

How Much Can a Few Engine Moves Help? Quantifying Limited Cheating in Chess

2026-01-08 · Daniel Keren arxiv

Cheating in chess, by using advice from powerful software, has become a major problem, reaching the highest levels. As opposed to the large majority of previous work, which concerned {\em detection} of cheating, here we …

Hyperparameter Optimization

Skewed Score: A statistical framework to assess autograders

2025-07-04 · Magda Dubois, Harry Coppock, Mario Giulianelli, Timo Flesch 외 arxiv

The evaluation of large language model (LLM) outputs is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation…

Bias Detection

Evaluating Complex Task through Crowdsourcing: Multiple Views Approach

2017-03-30 · Lingyu Lyu, Mehmed Kantardzic

With the popularity of massive open online courses, grading through crowdsourcing has become a prevalent approach towards large scale classes. However, for getting grades for complex tasks, which require specific skills …