Impacts of Histories and Models on LLM Grading: A Study in Advanced Software Engineering Courses
Graduate-level research reading report assessment creates a substantial labor burden for educators. While large language models (LLMs) hold great potential for automating academic grading, their reliability for this specialized task remains understudied, particularly regarding grading consistency, the lack of which represents a primary obstacle to educational fairness. This paper proposes a human-aligned LLM-assisted grading workflow and presents a case study based on 180 student submissions from a graduate advanced software engineering course. We evaluate two mainstream LLMs, Grok and GPT, in terms of grading consistency and alignment with human scores. We find LLMs exhibit distinct levels of intra-model consistency and significant inter-model grading inconsistencies, while simple ensemble approaches cannot improve alignment with human evaluation. Critically, continuous interaction history drives systematic drift in models' grading standards away from human expert scores. Our findings demonstrate LLMs' potential in reducing grading workload for educators in graduate education, while highlighting that indiscriminate LLM grading may introduce systemic unfairness, suggesting that specific operational practices are required to mitigate such disparities.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Learning and Upgrading in Global Value Chains: An Analysis of India's Manufacturing Sector
The topic of my research is "Learning and Upgrading in Global Value Chains: An Analysis of India's Manufacturing Sector". To analyse India's learning and upgrading through position, functions, specialisation & value addi…
PositionARAGOG: Advanced RAG Output Grading
Retrieval-Augmented Generation (RAG) is essential for integrating external knowledge into Large Language Model (LLM) outputs. While the literature on RAG is growing, it primarily focuses on systematic reviews and compari…
Document EmbeddingLanguage ModelingLanguage ModellingLarge Language Model+5An Analysis of ISO 26262: Using Machine Learning Safely in Automotive Software
Machine learning (ML) plays an ever-increasing role in advanced automotive functionality for driver assistance and autonomous operation; however, its adequacy from the perspective of safety certification remains controve…
BIG-bench Machine LearningSpineNetV2: Automated Detection, Labelling and Radiological Grading Of Clinical MR Scans
This technical report presents SpineNetV2, an automated tool which: (i) detects and labels vertebral bodies in clinical spinal magnetic resonance (MR) scans across a range of commonly used sequences; and (ii) performs ra…
Body DetectionAn Empirical Investigation into Learning Bug-Fixing Patches in the Wild via Neural Machine Translation
Millions of open-source projects with numerous bug fixes are available in code repositories. This proliferation of software development histories can be leveraged to learn how to fix common programming bugs. To explor…
Bug fixingDecoderMachine TranslationTranslation