paper-with-me

홈 › Papers

Synthetic Student Responses: LLM-Extracted Features for IRT Difficulty Parameter Estimation

2026-01-18 · Matias Hoyl arxiv

Educational assessment relies heavily on knowing question difficulty, traditionally determined through resource-intensive pre-testing with students. This creates significant barriers for both classroom teachers and assessment developers. We investigate whether Item Response Theory (IRT) difficulty parameters can be accurately estimated without student testing by modeling the response process and explore the relative contribution of different feature types to prediction accuracy. Our approach combines traditional linguistic features with pedagogical insights extracted using Large Language Models (LLMs), including solution step count, cognitive complexity, and potential misconceptions. We implement a two-stage process: first training a neural network to predict how students would respond to questions, then deriving difficulty parameters from these simulated response patterns. Using a dataset of over 250,000 student responses to mathematics questions, our model achieves a Pearson correlation of approximately 0.78 between predicted and actual difficulty parameters on completely unseen questions.

📄 PDF Abstract BibTeX arXiv:2602.00034

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction

2025-07-07 · Alexander Scarlatos, Nigel Fernandez, Christopher Ormerod, Susan Lottridge 외

Item (question) difficulties play a crucial role in educational assessments, enabling accurate and efficient assessment of student abilities and personalization to maximize learning outcomes. Traditionally, estimating it…

Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models

2026-01-05 · Christopher Ormerod arxiv

Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibratio…

Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading Efficiency

2023-10-10 · Eric Zelikman, Wanjing Anya Ma, Jasmine E. Tran, Diyi Yang 외

Developing an educational test can be expensive and time-consuming, as each item must be written by experts and then evaluated by collecting hundreds of student responses. Moreover, many tests require multiple distinct s…

Language ModelingLanguage ModellingSentence

Clustering students' open-ended questionnaire answers

2018-09-19 · Hämäläinen Wilhelmiina, Joy Mike, Berger Florian, Huttunen Sami

Open responses form a rich but underused source of information in educational data mining and intelligent tutoring systems. One of the major obstacles is the difficulty of clustering short texts automatically. In this pa…

Clustering

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

2026-06-17 · Han Chen, Ming Li, Chenguang Wang, Yijun Liang 외 arxiv

Item discrimination is a fundamental psychometric property of educational assessment, which measures whether an item meaningfully distinguishes students with higher proficiency from students with lower proficiency. While…

Reading Comprehension