paper-with-me

홈 › Papers

Skewed Score: A statistical framework to assess autograders

2025-07-04 · Magda Dubois, Harry Coppock, Mario Giulianelli, Timo Flesch, Lennart Luettgau, Cozmin Ududec arxiv

The evaluation of large language model (LLM) outputs is increasingly performed by other LLMs, a setup commonly known as "LLM-as-a-judge", or autograders. While autograders offer a scalable alternative to human evaluation, they have shown mixed reliability and may exhibit systematic biases, depending on response type, scoring methodology, domain specificity, or other factors. Here we propose a statistical framework based on Bayesian generalised linear models (GLMs) that enables researchers to simultaneously assess their autograders while addressing their primary research questions (e.g., LLM evaluation). Our approach models evaluation outcomes (e.g., scores or pairwise preferences) as a function of properties of the grader (e.g., human vs. autograder) and the evaluated item (e.g., response length or the LLM that generated it), allowing for explicit quantification of scoring differences and potential biases within a unified framework. In addition, our method can be used to augment traditional metrics such as inter-rater agreement, by providing uncertainty estimates and clarifying sources of disagreement. Overall, this approach contributes to more robust and interpretable use of autograders in LLM evaluation, enabling both performance analysis and bias detection.

📄 PDF Abstract BibTeX arXiv:2507.03772

Code (0)

등록된 구현이 없습니다.

Tasks

Bias Detection

Similar Papers 제목 키워드 기반

Forecasting Probability Distributions of Financial Returns with Deep Neural Networks

2025-08-26 · Jakub Michańków arxiv

This study evaluates deep neural networks for forecasting probability distributions of financial returns. 1D convolutional neural networks (CNN) and Long Short-Term Memory (LSTM) architectures are used to forecast parame…

PRNet: A Progressive Regression Network for No-Reference User-Generated-Content Video Quality Assessment

2021-09-29 · Yang YangR, Bo Jiang, Kailin Wu

Non-professional video, commonly known as User Generated Content (UGC) has become very popular in today’s video sharing applications. However, objectively perceptual quality assessment of UGC-videos is still a challenge …

regressionVideo Quality AssessmentVisual Question Answering (VQA)

JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

2026-05-24 · Russell Yang, Ruishi Chen, Pierce Kelaita, Riya Ranjan 외 arxiv

Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs. Although both met…

Are we Estimating or Guesstimating Translation Quality?

2020-07-01 · ACL 2020 6 · Shuo Sun, Francisco Guzm{\'a}n, Lucia Specia

Recent advances in pre-trained multilingual language models lead to state-of-the-art results on the task of quality estimation (QE) for machine translation. A carefully engineered ensemble of such models won the QE share…

Machine TranslationTranslation

An Analysis of Programming Course Evaluations Before and After the Introduction of an Autograder

2021-10-28 · Gerhard Johann Hagerer, Laura Lahesoo, Miriam Anschütz, Stephan Krusche 외

Commonly, introductory programming courses in higher education institutions have hundreds of participating students eager to learn to program. The manual effort for reviewing the submitted source code and for providing f…