paper-with-me

Papers

Bridging Qualitative Rubrics and AI: A Binary Question Framework for Criterion-Referenced Grading in Engineering

2026-01-22 · Lili Chen, Winn Wing-Yiu Chow, Stella Peng, Bencheng Fan, Sachitha Bandara arxiv

PURPOSE OR GOAL: This study investigates how GenAI can be integrated with a criterion-referenced grading framework to improve the efficiency and quality of grading for mathematical assessments in engineering. It specifically explores the challenges demonstrators face with manual, model solution-based grading and how a GenAI-supported system can be designed to reliably identify student errors, provide high-quality feedback, and support human graders. The research also examines human graders' perceptions of the effectiveness of this GenAI-assisted approach. ACTUAL OR ANTICIPATED OUTCOMES: The study found that GenAI achieved an overall grading accuracy of 92.5%, comparable to two experienced human graders. The two researchers, who also served as subject demonstrators, perceived the GenAI as a helpful second reviewer that improved accuracy by catching small errors and provided more complete feedback than they could manually. A central outcome was the significant enhancement of formative feedback. However, they noted the GenAI tool is not yet reliable enough for autonomous use, especially with unconventional solutions. CONCLUSIONS/RECOMMENDATIONS/SUMMARY: This study demonstrates that GenAI, when paired with a structured, criterion-referenced framework using binary questions, can grade engineering mathematical assessments with an accuracy comparable to human experts. Its primary contribution is a novel methodological approach that embeds the generation of high-quality, scalable formative feedback directly into the assessment workflow. Future work should investigate student perceptions of GenAI grading and feedback.

📄 PDF Abstract BibTeX arXiv:2601.15626

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning

2026-05-22 · Xiaoyuan Li, Keqin Bao, Moxin Li, Yubo Ma 외 arxiv

Rubric-based rewards offer a promising way to extend reinforcement learning (RL) for large language models beyond tasks with automatically verifiable answers. However, scaling rubric-based RL remains challenging: existin…

Reinforcement LearningContinual PretrainingInstruction Following

Online Rubrics Elicitation from Pairwise Comparisons

2025-10-08 · MohammadHossein Rezaei, Robert Vacareanu, Zihao Wang, Clinton Wang 외 arxiv

Rubrics provide a flexible way to train LLMs on open-ended long-form answers where verifiable rewards are not applicable and human preferences provide coarse signals. Prior work shows that reinforcement learning with rub…

Reinforcement Learning

Qworld: Question-Specific Evaluation Criteria for LLMs

2026-03-06 · Shanghua Gao, Yuchang Su, Pengwei Sui, Curtis Ginder 외 arxiv

Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Binary scores and static rubrics fail to capture these context-dependent requiremen…

C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences

2026-04-15 · Akira Kawabata, Saku Sugawara arxiv

Rubric-augmented verification guides reward models with explicit evaluation criteria, yielding more reliable judgments than single-model verification. However, most existing methods require costly rubric annotations, lim…

Gemini Pro Defeated by GPT-4V: Evidence from Education

2023-12-27 · Gyeong-Geon Lee, Ehsan Latif, Lehong Shi, Xiaoming Zhai

This study compared the classification performance of Gemini Pro and GPT-4V in educational settings. Employing visual question answering (VQA) techniques, the study examined both models' abilities to read text-based rubr…

image-classificationImage ClassificationQuestion AnsweringVisual Question Answering+1