paper-with-me

Papers

Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction

2025-12-21 · Ming Li, Han Chen, Yunze Xiao, Jian Chen, Hong Jiao, Tianyi Zhou arxiv

Accurate estimation of item (question or task) difficulty is critical for educational assessment but suffers from the cold start problem. While Large Language Models demonstrate superhuman problem-solving capabilities, it remains an open question whether they can perceive the cognitive struggles of human learners. In this work, we present a large-scale empirical analysis of Human-AI Difficulty Alignment for over 20 models across diverse domains such as medical knowledge and mathematical reasoning. Our findings reveal a systematic misalignment where scaling up model size is not reliably helpful; instead of aligning with humans, models converge toward a shared machine consensus. We observe that high performance often impedes accurate difficulty estimation, as models struggle to simulate the capability limitations of students even when being explicitly prompted to adopt specific proficiency levels. Furthermore, we identify a critical lack of introspection, as models fail to predict their own limitations. These results suggest that general problem-solving capability does not imply an understanding of human cognitive struggles, highlighting the challenge of using current models for automated difficulty prediction.

📄 PDF Abstract BibTeX arXiv:2512.18880

Code (0)

등록된 구현이 없습니다.

Tasks

Mathematical Reasoning

Similar Papers 제목 키워드 기반

Interpretable Difficulty-Aware Knowledge Tracing in Tutor-Student Dialogues

2026-05-01 · Shuyan Huang, Alexander Scarlatos, Jaewook Lee, Andrew Lan arxiv

Recent advances in large language models (LLMs) have led to the development of AI-powered tutoring systems that provide interactive support via dialogue. To enable these tutoring systems to provide personalized support, …

Knowledge Tracing

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

2026-06-17 · Han Chen, Ming Li, Chenguang Wang, Yijun Liang 외 arxiv

Item discrimination is a fundamental psychometric property of educational assessment, which measures whether an item meaningfully distinguishes students with higher proficiency from students with lower proficiency. While…

Reading Comprehension

Do LLMs Implicitly Determine the Suitable Text Difficulty for Users?

2024-02-22 · Seiji Gobara, Hidetaka Kamigaito, Taro Watanabe

Education that suits the individual learning level is necessary to improve students' understanding. The first step in achieving this purpose by using large language models (LLMs) is to adjust the textual difficulty of th…

Question Answering

Take Out Your Calculators: Estimating the Real Difficulty of Question Items with LLM Student Simulations

2026-01-15 · Christabel Acquaye, Yi Ting Huang, Marine Carpuat, Rachel Rudinger arxiv

Standardized math assessments require expensive human pilot studies to establish the difficulty of test items. We investigate the predictive value of open-source large language models (LLMs) for evaluating the difficulty…

Synthetic Student Responses: LLM-Extracted Features for IRT Difficulty Parameter Estimation

2026-01-18 · Matias Hoyl arxiv

Educational assessment relies heavily on knowing question difficulty, traditionally determined through resource-intensive pre-testing with students. This creates significant barriers for both classroom teachers and asses…