paper-with-me

홈 › Papers

Impacts of Histories and Models on LLM Grading: A Study in Advanced Software Engineering Courses

2026-06-07 · Qilin Zhou, Zhuo Wang, Yue Li, W. K. Chan arxiv

Graduate-level research reading report assessment creates a substantial labor burden for educators. While large language models (LLMs) hold great potential for automating academic grading, their reliability for this specialized task remains understudied, particularly regarding grading consistency, the lack of which represents a primary obstacle to educational fairness. This paper proposes a human-aligned LLM-assisted grading workflow and presents a case study based on 180 student submissions from a graduate advanced software engineering course. We evaluate two mainstream LLMs, Grok and GPT, in terms of grading consistency and alignment with human scores. We find LLMs exhibit distinct levels of intra-model consistency and significant inter-model grading inconsistencies, while simple ensemble approaches cannot improve alignment with human evaluation. Critically, continuous interaction history drives systematic drift in models' grading standards away from human expert scores. Our findings demonstrate LLMs' potential in reducing grading workload for educators in graduate education, while highlighting that indiscriminate LLM grading may introduce systemic unfairness, suggesting that specific operational practices are required to mitigate such disparities.

📄 PDF Abstract BibTeX arXiv:2606.08400

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning and Upgrading in Global Value Chains: An Analysis of India's Manufacturing Sector

2021-01-12 · Sourish Dutta

The topic of my research is "Learning and Upgrading in Global Value Chains: An Analysis of India's Manufacturing Sector". To analyse India's learning and upgrading through position, functions, specialisation & value addi…

Position

ARAGOG: Advanced RAG Output Grading

2024-04-01 · Matouš Eibich, Shivay Nagpal, Alexander Fred-Ojala

Retrieval-Augmented Generation (RAG) is essential for integrating external knowledge into Large Language Model (LLM) outputs. While the literature on RAG is growing, it primarily focuses on systematic reviews and compari…

Document EmbeddingLanguage ModelingLanguage ModellingLarge Language Model+5

An Analysis of ISO 26262: Using Machine Learning Safely in Automotive Software

2017-09-07 · Rick Salay, Rodrigo Queiroz, Krzysztof Czarnecki

Machine learning (ML) plays an ever-increasing role in advanced automotive functionality for driver assistance and autonomous operation; however, its adequacy from the perspective of safety certification remains controve…

BIG-bench Machine Learning

SpineNetV2: Automated Detection, Labelling and Radiological Grading Of Clinical MR Scans

2022-05-03 · Rhydian Windsor, Amir Jamaludin, Timor Kadir, Andrew Zisserman

This technical report presents SpineNetV2, an automated tool which: (i) detects and labels vertebral bodies in clinical spinal magnetic resonance (MR) scans across a range of commonly used sequences; and (ii) performs ra…

Body Detection

An Empirical Investigation into Learning Bug-Fixing Patches in the Wild via Neural Machine Translation

2018-09-07 · Accepted to the ACM Transactions on Software Engineering and Methodology 2018 9 · Michele Tufano College of William and Mary Williamsburg, USA Cody Watson College of William and Mary Williamsburg, USA Gabriele Bavota Università della Svizzera italiana (USI) Lugano, Switzerland Massimiliano Di Penta University of Sannio Benevento 외

Millions of open-source projects with numerous bug fixes are available in code repositories. This proliferation of software development histories can be leveraged to learn how to fix common programming bugs. To explor…

Bug fixingDecoderMachine TranslationTranslation