paper-with-me

홈 › Papers

A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization

2024-10-03 · Yucheng Chu, Hang Li, Kaiqi Yang, Harry Shomer, Hui Liu, Yasemin Copur-Gencturk, Jiliang Tang

Open-ended short-answer questions (SAGs) have been widely recognized as a powerful tool for providing deeper insights into learners' responses in the context of learning analytics (LA). However, SAGs often present challenges in practice due to the high grading workload and concerns about inconsistent assessments. With recent advancements in natural language processing (NLP), automatic short-answer grading (ASAG) offers a promising solution to these challenges. Despite this, current ASAG algorithms are often limited in generalizability and tend to be tailored to specific questions. In this paper, we propose a unified multi-agent ASAG framework, GradeOpt, which leverages large language models (LLMs) as graders for SAGs. More importantly, GradeOpt incorporates two additional LLM-based agents - the reflector and the refiner - into the multi-agent system. This enables GradeOpt to automatically optimize the original grading guidelines by performing self-reflection on its errors. Through experiments on a challenging ASAG task, namely the grading of pedagogical content knowledge (PCK) and content knowledge (CK) questions, GradeOpt demonstrates superior performance in grading accuracy and behavior alignment with human graders compared to representative baselines. Finally, comprehensive ablation studies confirm the effectiveness of the individual components designed in GradeOpt.

📄 PDF Abstract BibTeX arXiv:2410.02165

Code (0)

등록된 구현이 없습니다.

Tasks

automatic short answer grading

Similar Papers 제목 키워드 기반

LLM-based Automated Grading with Human-in-the-Loop

2025-04-07 · Hang Li, Yucheng Chu, Kaiqi Yang, Yasemin Copur-Gencturk 외

The rise of artificial intelligence (AI) technologies, particularly large language models (LLMs), has brought significant advancements to the field of education. Among various applications, automatic short answer grading…

automatic short answer grading

Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

2026-08-27 · Mingqi Gao, Anthony Sicilia, Weiyan Shi arxiv

Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered e…

Auditing an Automatic Grading Model with deep Reinforcement Learning

2024-05-11 · Aubrey Condor, Zachary Pardos

We explore the use of deep reinforcement learning to audit an automatic short answer grading (ASAG) model. Automatic grading may decrease the time burden of rating open-ended items for educators, but a lack of robust eva…

automatic short answer gradingDeep Reinforcement Learningreinforcement-learningReinforcement Learning

CHiL(L)Grader: Calibrated Human-in-the-Loop Short-Answer Grading

2026-03-12 · Pranav Raikote, Korbinian Randl, Ioanna Miliou, Athanasios Lakes 외 arxiv

Scaling educational assessment with large language models requires not just accuracy, but the ability to recognize when predictions are trustworthy. Instruction-tuned models tend to be overconfident, and their reliabilit…

Continual Learning

Human and Automated CEFR-based Grading of Short Answers

2017-09-01 · WS 2017 9 · Ana{\"\i}s Tack, Thomas Fran{\c{c}}ois, Sophie Roekhaut, C{\'e}drick Fairon

This paper is concerned with the task of automatically assessing the written proficiency level of non-native (L2) learners of English. Drawing on previous research on automated L2 writing assessment following the Common …