Predicting Item Survival for Multiple Choice Questions in a High-Stakes Medical Exam
One of the most resource-intensive problems in the educational testing industry relates to ensuring that newly-developed exam questions can adequately distinguish between students of high and low ability. The current practice for obtaining this information is the costly procedure of pretesting: new items are administered to test-takers and then the items that are too easy or too difficult are discarded. This paper presents the first study towards automatic prediction of an item{'}s probability to {``}survive{''} pretesting (item survival), focusing on human-produced MCQs for a medical exam. Survival is modeled through a number of linguistic features and embedding types, as well as features inspired by information retrieval. The approach shows promising first results for this challenging new application and for modeling the difficulty of expert-knowledge questions.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalMultiple-choiceRetrievalSimilar Papers 제목 키워드 기반
Predicting the Difficulty and Response Time of Multiple Choice Questions Using Transfer Learning
This paper investigates whether transfer learning can improve the prediction of the difficulty and response time parameters for 18,000 multiple-choice questions from a high-stakes medical exam. The type the signal that b…
Multiple-choiceTransfer LearningUnibucLLM: Harnessing LLMs for Automated Prediction of Item Difficulty and Response Time for Multiple-Choice Questions
This work explores a novel data augmentation method based on Large Language Models (LLMs) for predicting item difficulty and response time of retired USMLE Multiple-Choice Questions (MCQs) in the BEA 2024 Shared Task. Ou…
Data AugmentationMultiple-choicePredicting the Difficulty of Multiple Choice Questions in a High-stakes Medical Exam
Predicting the construct-relevant difficulty of Multiple-Choice Questions (MCQs) has the potential to reduce cost while maintaining the quality of high-stakes exams. In this paper, we propose a method for estimating the …
Multiple-choiceQuestion AnsweringAssessing the Quality of Multiple-Choice Questions Using GPT-4 and Rule-Based Methods
Multiple-choice questions with item-writing flaws can negatively impact student learning and skew analytics. These flaws are often present in student-generated questions, making it difficult to assess their quality and s…
Multiple-choiceTell Me Who Your Students Are: GPT Can Generate Valid Multiple-Choice Questions When Students' (Mis)Understanding Is Hinted
The primary goal of this study is to develop and evaluate an innovative prompting technique, AnaQuest, for generating multiple-choice questions (MCQs) using a pre-trained large language model. In AnaQuest, the choice ite…
Language ModelingLanguage ModellingLarge Language ModelMultiple-choice+2