Using Large Language Models to Assess Teachers' Pedagogical Content Knowledge
Assessing teachers' pedagogical content knowledge (PCK) through performance-based tasks is both time and effort-consuming. While large language models (LLMs) offer new opportunities for efficient automatic scoring, little is known about whether LLMs introduce construct-irrelevant variance (CIV) in ways similar to or different from traditional machine learning (ML) and human raters. This study examines three sources of CIV -- scenario variability, rater severity, and rater sensitivity to scenario -- in the context of video-based constructed-response tasks targeting two PCK sub-constructs: analyzing student thinking and evaluating teacher responsiveness. Using generalized linear mixed models (GLMMs), we compared variance components and rater-level scoring patterns across three scoring sources: human raters, supervised ML, and LLM. Results indicate that scenario-level variance was minimal across tasks, while rater-related factors contributed substantially to CIV, especially in the more interpretive Task II. The ML model was the most severe and least sensitive rater, whereas the LLM was the most lenient. These findings suggest that the LLM contributes to scoring efficiency while also introducing CIV as human raters do, yet with varying levels of contribution compared to supervised ML. Implications for rater training, automated scoring design, and future research on model interpretability are discussed.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
How Useful are Educational Questions Generated by Large Language Models?
Controllable text generation (CTG) by large language models has a huge potential to transform education for teachers and students alike. Specifically, high quality and diverse question generation can dramatically reduce …
Question GenerationQuestion-GenerationText GenerationTo what extent is ChatGPT useful for language teacher lesson plan creation?
The advent of generative AI models holds tremendous potential for aiding teachers in the generation of pedagogical materials. However, numerous knowledge gaps concerning the behavior of these models obfuscate the generat…
SpecificityDesign and Evaluation for a Prototype of an Online Tool to Access Mathematics Notions in Sign Language
The Sign{'}Maths project aims at giving access to pedagogical resources in Sign Language (SL). It will provide Deaf students and teachers with mathematics vocabulary in SL, this in order to contribute to the standardisat…
NavigateEnabling Multi-Agent Systems as Learning Designers: Applying Learning Sciences to AI Instructional Design
K-12 educators are increasingly using Large Language Models (LLMs) to create instructional materials. These systems excel at producing fluent, coherent content, but often lack support for high-quality teaching. The reaso…
Prompt EngineeringHow Teachers Can Use Large Language Models and Bloom's Taxonomy to Create Educational Quizzes
Question generation (QG) is a natural language processing task with an abundance of potential benefits and use cases in the educational domain. In order for this potential to be realized, QG systems must be designed and …
Language ModelingLanguage ModellingLarge Language ModelQuestion Generation+1