paper-with-me

홈 › Papers

Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory

2026-04-30 · Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Hendrik Drachsler, Ulf Kroehne arxiv

Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa. However, these metrics provide limited insight into how grading performance varies across student responses of differing grading difficulty. We introduce an evaluation framework for LLM-based ASAG based on item response theory (IRT), which models grading correctness as a function of latent grader ability and response grading difficulty. This formulation enables response-level analysis of where LLM graders succeed or fail and reveals robustness differences that are not visible from aggregate scores alone. We apply the framework to 17 open-weight LLMs on the SciEntsBank and Beetle benchmarks. The results show that even models with similar overall performance differ substantially in how sharply their grading accuracy declines as response difficulty increases. In addition, confusion patterns show that errors on difficult responses concentrate disproportionately on the \texttt{partially\_correct\_incomplete} label, indicating a tendency toward intermediate-label collapse under ambiguity. To characterize difficult responses, we further analyze semantic and linguistic correlates of estimated difficulty. Across both datasets, higher difficulty is associated with weaker semantic alignment to the reference answer, stronger contradiction signals, and greater semantic isolation in embedding space. Overall, these results show that item response theory offers a useful framework for evaluating LLM-based ASAG beyond aggregate performance measures.

📄 PDF Abstract BibTeX arXiv:2605.00238

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Auditing an Automatic Grading Model with deep Reinforcement Learning

2024-05-11 · Aubrey Condor, Zachary Pardos

We explore the use of deep reinforcement learning to audit an automatic short answer grading (ASAG) model. Automatic grading may decrease the time burden of rating open-ended items for educators, but a lack of robust eva…

automatic short answer gradingDeep Reinforcement Learningreinforcement-learningReinforcement Learning

Cognitively Aided Zero-Shot Automatic Essay Grading

2021-02-22 · ICON 2020 12 · Sandeep Mathias, Rudra Murthy, Diptesh Kanojia, Pushpak Bhattacharyya

Automatic essay grading (AEG) is a process in which machines assign a grade to an essay written in response to a topic, called the prompt. Zero-shot AEG is when we train a system to grade essays written to a new prompt w…

Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams

2024-11-07 · Adriana Caraeni, Alexander Scarlatos, Andrew Lan

Recent advances in generative artificial intelligence (AI) have shown promise in accurately grading open-ended student responses. However, few prior works have explored grading handwritten responses due to a lack of data…

Math

Learning Robust Representation for Joint Grading of Ophthalmic Diseases via Adaptive Curriculum and Feature Disentanglement

2022-07-09 · Haoxuan Che, Haibo Jin, Hao Chen

Diabetic retinopathy (DR) and diabetic macular edema (DME) are leading causes of permanent blindness worldwide. Designing an automatic grading system with good generalization ability for DR and DME is vital in clinical p…

Disentanglement

Powergrading: a Clustering Approach to Amplify Human Effort for Short Answer Grading

2013-01-01 · TACL 2013 1 · Sumit Basu, Chuck Jacobs, V, Lucy erwende

We introduce a new approach to the machine-assisted grading of short answer questions. We follow past work in automated grading by first training a similarity metric between student responses, but then go on to use this …

Clustering