paper-with-me

Papers

ChatGPT for automated grading of short answer questions in mechanical ventilation

2025-05-05 · Tejas Jade, Alex Yartsev

Standardised tests using short answer questions (SAQs) are common in postgraduate education. Large language models (LLMs) simulate conversational language and interpret unstructured free-text responses in ways aligning with applying SAQ grading rubrics, making them attractive for automated grading. We evaluated ChatGPT 4o to grade SAQs in a postgraduate medical setting using data from 215 students (557 short-answer responses) enrolled in an online course on mechanical ventilation (2020--2024). Deidentified responses to three case-based scenarios were presented to ChatGPT with a standardised grading prompt and rubric. Outputs were analysed using mixed-effects modelling, variance component analysis, intraclass correlation coefficients (ICCs), Cohen's kappa, Kendall's W, and Bland--Altman statistics. ChatGPT awarded systematically lower marks than human graders with a mean difference (bias) of -1.34 on a 10-point scale. ICC values indicated poor individual-level agreement (ICC1 = 0.086), and Cohen's kappa (-0.0786) suggested no meaningful agreement. Variance component analysis showed minimal variability among the five ChatGPT sessions (G-value = 0.87), indicating internal consistency but divergence from the human grader. The poorest agreement was observed for evaluative and analytic items, whereas checklist and prescriptive rubric items had less disagreement. We caution against the use of LLMs in grading postgraduate coursework. Over 60% of ChatGPT-assigned grades differed from human grades by more than acceptable boundaries for high-stakes assessments.

📄 PDF Abstract BibTeX arXiv:2505.04645

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Grading Conversational Responses Of Chatbots

2023-02-01 · Grant Rosario, David Noever

Chatbots have long been capable of answering basic questions and even responding to obscure prompts, but recently their improvements have been far more significant. Modern chatbots like Open AIs ChatGPT3 not only have th…

Machine TranslationTranslation

ASAG2024: A Combined Benchmark for Short Answer Grading

2024-09-27 · Gérôme Meyer, Philip Breuer, Jonathan Fürst

Open-ended questions test a more thorough understanding than closed-ended questions and are often a preferred assessment method. However, open-ended questions are tedious to grade and subject to personal bias. Therefore,…

"Did my figure do justice to the answer?" : Towards Multimodal Short Answer Grading with Feedback (MMSAF)

2024-12-27 · Pritam Sil, Pushpak Bhattacharyya

Assessments play a vital role in a student's learning process. This is because they provide valuable feedback crucial to a student's growth. Such assessments contain questions with open-ended responses, which are difficu…

automatic short answer grading

Powergrading: a Clustering Approach to Amplify Human Effort for Short Answer Grading

2013-01-01 · TACL 2013 1 · Sumit Basu, Chuck Jacobs, V, Lucy erwende

We introduce a new approach to the machine-assisted grading of short answer questions. We follow past work in automated grading by first training a similarity metric between student responses, but then go on to use this …

Clustering

Human and Automated CEFR-based Grading of Short Answers

2017-09-01 · WS 2017 9 · Ana{\"\i}s Tack, Thomas Fran{\c{c}}ois, Sophie Roekhaut, C{\'e}drick Fairon

This paper is concerned with the task of automatically assessing the written proficiency level of non-native (L2) learners of English. Drawing on previous research on automated L2 writing assessment following the Common …