paper-with-me

홈 › Papers

Confidence Estimation in Automatic Short Answer Grading with LLMs

2026-04-30 · Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Hendrik Drachsler, Ulf Kroehne arxiv

Automatic Short Answer Grading (ASAG) with generative large language models (LLMs) has recently demonstrated strong performance without task-specific fine-tuning, while also enabling the generation of synthetic feedback for educational assessment. Despite these advances, LLM-based grading remains imperfect, making reliable confidence estimates essential for safe and effective human-AI collaboration in educational decision-making. In this work, we investigate confidence estimation for ASAG with LLMs by jointly considering model-based confidence signals and dataset-derived uncertainty. We systematically compare three model-based confidence estimation strategies, namely verbalizing, latent, and consistency-based confidence estimation, and show that model-based confidence alone is insufficient to reliably capture uncertainty in ASAG. To address this limitation, we propose a hybrid confidence framework that integrates model-based confidence signals with an explicit estimate of dataset-derived aleatoric uncertainty. Aleatoric uncertainty is operationalized by clustering semantically embedded student responses and quantifying within-cluster heterogeneity. Our results demonstrate that the proposed hybrid confidence measure yields more reliable confidence estimates and improves selective grading performance compared to single-source approaches. Overall, this work advances confidence-aware LLM-based grading for human-in-the-loop assessment, supporting more trustworthy AI-assisted educational assessment systems.

📄 PDF Abstract BibTeX arXiv:2605.00200

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading

2026-03-12 · Pranav Raikote, Korbinian Randl, Ioanna Miliou, Athanasios Lakes 외 arxiv

Scaling educational assessment with large language models requires not just accuracy, but the ability to recognize when predictions are trustworthy. Instruction-tuned models tend to be overconfident, and their reliabilit…

Continual Learning

Balancing Cost and Quality: An Exploration of Human-in-the-loop Frameworks for Automated Short Answer Scoring

2022-06-16 · Hiroaki Funayama, Tasuku Sato, Yuichiroh Matsubayashi, Tomoya Mizumoto 외

Short answer scoring (SAS) is the task of grading short text written by a learner. In recent years, deep-learning-based approaches have substantially improved the performance of SAS models, but how to guarantee high-qual…

When Can We Trust LLM Graders? Calibrating Confidence for Automated Assessment

2026-03-31 · Robinson Ferrer, Damla Turgut, Zhongzhou Chen, Shashank Sonkar arxiv

Large Language Models (LLMs) show promise for automated grading, but their outputs can be unreliable. Rather than improving grading accuracy directly, we address a complementary problem: \textit{predicting when an LLM gr…

Cheating Automatic Short Answer Grading: On the Adversarial Usage of Adjectives and Adverbs

2022-01-20 · Anna Filighera, Sebastian Ochs, Tim Steuer, Thomas Tregel

Automatic grading models are valued for the time and effort saved during the instruction of large student bodies. Especially with the increasing digitization of education and interest in large-scale standardized testing,…

Adversarial Attackautomatic short answer gradingvalid

AR-ASAG An ARabic Dataset for Automatic Short Answer Grading Evaluation

2020-05-01 · LREC 2020 5 · Leila Ouahrani, Djamal Bennouar

Automatic short answer grading is a significant problem in E-assessment. Several models have been proposed to deal with it. Evaluation and comparison of such solutions need the availability of Datasets with manual exampl…

automatic short answer gradingSpecificity